Wan 2.2 API
Wan 2.2 của Alibaba — bộ công cụ video mã nguồn mở trọng số triển khai trên WaveSpeedAI với hơn 35 biến thể chính chủ: Animate (hoạt hình nhân vật 120 giây), Video Edit, Speech-to-Video (10 phút theo âm thanh), Fun-Control (giấy phép Apache 2.0), cùng image-to-video và text-to-video ở nhiều cỡ mô hình (5B, A14B) và độ phân giải (480p / 720p).
Chỉ gồm các biến thể do WaveSpeedAI lưu trữ. Animate tạo clip 720p dài tới 120 giây; Speech-to-Video tạo clip 480p dài tới 10 phút; Fun-Control dùng Control Code cài sẵn theo giấy phép Apache 2.0 cho mục đích thương mại. Endpoint huấn luyện LoRA tinh chỉnh chỉ trong vài phút.
Tổng quan
Giới thiệu về Wan 2.2 API
Wan 2.2 làm được gì, nằm ở đâu trong dòng mô hình của Alibaba, và vì sao các đội ngũ chọn nó.
Wan 2.2 là mô hình tạo video từ Alibaba, có sẵn qua REST API của WaveSpeedAI. Wan 2.2 của Alibaba — bộ công cụ video mã nguồn mở trọng số triển khai trên WaveSpeedAI với hơn 35 biến thể chính chủ: Animate (hoạt hình nhân vật 120 giây), Video Edit, Speech-to-Video (10 phút theo âm thanh), Fun-Control (giấy phép Apache 2.0), cùng image-to-video và text-to-video ở nhiều cỡ mô hình (5B, A14B) và độ phân giải (480p / 720p).
Chỉ gồm các biến thể do WaveSpeedAI lưu trữ. Animate tạo clip 720p dài tới 120 giây; Speech-to-Video tạo clip 480p dài tới 10 phút; Fun-Control dùng Control Code cài sẵn theo giấy phép Apache 2.0 cho mục đích thương mại. Endpoint huấn luyện LoRA tinh chỉnh chỉ trong vài phút.
Họ Wan 2.2 trên WaveSpeedAI cung cấp 32 endpoint REST bao gồm 8 quy trình Image-To-Video, Motion-Control, Video-To-Video, Image-To-Image, Digital-Human, Text-To-Image, Training, Text-To-Video. Mỗi biến thể có giá, các tham số điều chỉnh và kết quả mẫu riêng — hãy chọn biến thể phù hợp với dạng đầu vào và ràng buộc sản xuất của bạn, hoặc gọi nhiều biến thể từ cùng một API key để ghép các pipeline nhiều bước.
Chạy Wan 2.2 bằng cùng API key, tài khoản thanh toán và hạn mức tốc độ bạn dùng cho hơn 1.000 mô hình AI khác trên WaveSpeedAI. Không cần thiết lập nhà cung cấp riêng, không cần SDK cho từng nhà cung cấp, không có hạn mức tốc độ riêng của từng hãng — một lần tích hợp bao quát mọi thứ, từ text-to-image và text-to-video đến tổng hợp âm thanh, tạo 3D, nâng cấp và chỉnh sửa.
Endpoint
Tất cả endpoint API Wan 2.2
32 endpoint Wan 2.2 hiện có sẵn trên WaveSpeedAI — hãy chọn biến thể phù hợp với quy trình của bạn.
/filters:quality(82)/media/images/20260408111216_2fizlz7x.webp)
Wan 2.2 Image To Video Lora
Wan-2.2/image-to-video-lora enables unlimited image-to-video generation from a single image, producing smooth, cinematic motion with clean detail. Supports custom LoRAs for style and character consistency. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
/filters:quality(82)/media/images/20260408111152_zehbgma1.webp)
Wan 2.2 Image To Video
Wan 2.2 Image-to-Video turns a single image into smooth, cinematic motion with clean detail—ideal for storyboards, mood shots, and product demos. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
/filters:quality(82)/media/images/20260408111211_p8f6ffxe.webp)
Wan 2.2 Animate
Wan2.2-Animate unified character animation & replacement model replicating movement and expression; generates 720p videos up to 120s. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111241_3bzr5rd1.webp)
Wan 2.2 Video Edit
Wan 2.2 Video Edit lets you modify videos via text prompts (e.g., change clothing or characters). Powered by Wan 2.2, it supports 480p ($0.20/5s) and 720p ($0.40/5s), up to 120s. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111257_4rmbjsgj.webp)
Wan 2.2 Image To Image
WAN 2.2 (14B) is an image-to-image model for high-resolution photorealistic image editing with exceptional precision and fidelity. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111154_gvj4f5sj.webp)
Wan 2.2 Speech To Video
Wan-2.2-S2V turns images and speech into high-fidelity videos with realistic face and body motion; supports up to 10-minute clips in 480p, from $0.15/5s. Ready-to-use REST API, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111206_p4ked9pu.webp)
Wan 2.2 Text To Image Lora
WAN 2.2 generates super-detailed images from text prompts and supports custom LoRAs for fine-grained style and subject control. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111250_80jkk4zs.webp)
Wan 2.2 Fun Control
Wan2.2-Fun-Control uses Control Codes and multi-modal inputs to generate preset-controlled videos up to 120s at 720p; released under Apache 2.0 for commercial use. Ready-to-use REST API, no coldstarts, affordable.
/filters:quality(82)/media/images/20260408112009_zkynaf2n.webp)
Wan 2.2 Image Lora Trainer
Train custom Wan 2.2 character/style LoRA models 10x faster. Style training, character training, object training. From concept to model in minutes, not hours. Upload a ZIP file containing images to start!
/filters:quality(82)/media/images/20260408111954_4jny6q8q.webp)
Wan 2.2 I2v Lora Trainer
Train custom Wan 2.2 I2V LoRA models 10x faster. Action training, motion training, video efect training. From concept to model in minutes, not hours. Upload a ZIP file containing videos to start!
/filters:quality(82)/media/images/20260408111226_b3370qq2.webp)
Wan 2.2 Text To Image Realism
WAN 2.2 delivers ultra-realistic text-to-image generation, converting prompts into photoreal images with high fidelity and detail. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111246_kbtz9m3g.webp)
Wan 2.2 I2v 5b 720p Lora
Wan 2.2 i2v-5B-720p is a 5B image-to-video model producing 720p videos with LoRA support for style customization. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111210_znuyn6r8.webp)
Wan 2.2 T2v 5b 720p Lora
Wan 2.2 T2V 5B is a 5B text-to-video model with LoRA support that generates 720p videos from text prompts for easy personalization. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111236_0was1j4o.webp)
Wan 2.2 I2v 480p Lora Ultra Fast
Wan 2.2 i2v delivers ultra-fast Image-to-Video at 480p with support for custom LoRAs for tailored styles. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111252_sejz6k48.webp)
Wan 2.2 I2v 480p Ultra Fast
Wan 2.2 A14B Image-to-Video (i2v-480p) produces ultra-fast 480p videos from single images, enabling unlimited AI video generation with high throughput. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111149_z1lz7r9s.webp)
Wan 2.2 I2v 720p Ultra Fast
Generate unlimited ultra-fast 720p AI videos from images with Wan 2.2 A14B image-to-video model. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111259_8bca6477.webp)
Wan 2.2 T2v 480p Lora Ultra Fast
Ultra-fast Wan 2.2 text-to-video model producing 480p videos with custom LoRA support—generate unlimited AI videos with personalized styles. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111238_h7qxvir0.webp)
Wan 2.2 I2v 720p Lora Ultra Fast
Wan 2.2 i2v 720P is an ultra-fast Image-to-Video model that generates unlimited AI videos and supports custom LoRAs for personalized outputs. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111248_iyye47z8.webp)
Wan 2.2 T2v 480p Ultra Fast
Wan 2.2 t2v 480p Ultra-Fast generates unlimited AI videos from text prompts at 480p with ultra-fast inference. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111201_lv1td92d.webp)
Wan 2.2 T2v 5b 720p
Wan 2.2 T2V 5B is a 720P text-to-video model that generates unlimited AI videos from simple text prompts, producing consistent high-quality 720p outputs. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111221_xo8hpiqi.webp)
Wan 2.2 I2v 5b 720p
Wan 2.2 I2V 5B converts images into high-quality 720P videos using a 5B image-to-video model for AI video generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111214_rct2hj0c.webp)
Wan 2.2 I2v 480p
Wan 2.2 A14B converts images into 480p videos, enabling unlimited AI video generation from single images. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111234_cplw518o.webp)
Wan 2.2 I2v 480p Lora
WAN 2.2 A14B Image-to-Video model generates unlimited 480p videos from images and supports custom LoRAs for personalized styles and fine-tuning. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111232_8tgp9v7m.webp)
Wan 2.2 I2v 720p Lora
WAN 2.2 Image-to-Video (i2v) 720p converts images into 720p videos and supports custom LoRAs for style personalization. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111239_qlur3dn5.webp)
Wan 2.2 I2v 720p
WAN 2.2 A14B i2v-720p converts images into smooth 720p videos, enabling unlimited AI video generation with the Wan 2.2 image-to-video model. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111156_b3olejjy.webp)
Wan 2.2 T2v 480p
Wan 2.2 t2v-480p generates unlimited AI videos from text prompts at 480p resolution, ideal for rapid prototyping and content creation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111208_sw0j3xam.webp)
Wan 2.2 T2v 480p Lora
WAN 2.2 T2V 480p with LoRA generates text-to-video at 480p and supports custom LoRAs for personalized styles. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111218_00vzkvbl.webp)
Wan 2.2 T2v 720p
Wan 2.2 t2v-720p converts text prompts into native 720P videos, producing high-quality 720P clips from simple prompts. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111230_mtxgxh1x.webp)
Wan 2.2 T2v 720p Lora
Wan 2.2 T2V 720p with custom LoRA support turns text prompts into 720p AI videos and enables unlimited video generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111158_csf9651o.webp)
Wan 2.2 T2v 720p Lora Ultra Fast
Ultra-fast Wan 2.2 Text-to-Video generates unlimited 720p AI videos with custom LoRAs for personalized styles. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/20260408111244_aqp0x6b0.webp)
Wan 2.2 T2v 720p Ultra Fast
WAN 2.2 T2V 720p Ultra-Fast generates high-quality 720p videos from text prompts with unlimited output and ultra-fast throughput. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/1786269272686834557_4oOX6gqz.webp)
Wan 2.2 Animate 2
Wan 2.2 Animate 2 is the next-generation Wan character animation model: an end-to-end DiT that makes the character in a reference image perform the motion of a driving video, with no pose extraction, prompt-controlled background, and strong identity preservation; generates 480p/720p videos up to 120s. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Ví dụ
Xem Wan 2.2 hoạt động
Kết quả thật do API Wan 2.2 tạo ra. Rê chuột lên video để xem trước, nhấp để mở trình xem kích thước đầy đủ.
Hướng dẫn
Cách dùng API Wan 2.2
Bốn bước từ đăng ký đến khi có kết quả. Ví dụ đầy đủ bằng Python, Node.js và cURL nằm ở phần API bên dưới.
- 01
Lấy API key
Đăng ký tài khoản WaveSpeedAI và sao chép API key từ bảng điều khiển. Tài khoản mới được tặng credit dùng thử miễn phí — đủ để chạy playground vài chục lần trước khi bắt đầu tính phí.
- 02
Gửi một prediction
POST đầu vào của bạn dưới dạng JSON tới https://api.wavespeed.ai/api/v3/wavespeed-ai/wan-2.2/image-to-video. Endpoint trả về prediction id ngay lập tức — việc tạo là bất đồng bộ nên bạn không phải giữ kết nối mở trong lúc suy luận.
- 03
Poll để chờ hoàn tất
GET https://api.wavespeed.ai/api/v3/predictions/{request_id}/result. Trả về kết quả khi completed; dừng và báo lỗi khi failed, cancelled, timeout hoặc deleted; tiếp tục poll với mọi trạng thái khác.
- 04
Đọc URL kết quả
Khi status là "completed", đọc URL từ data.outputs[0]. URL trỏ tới media bạn đã tạo trên CDN của WaveSpeedAI — ảnh, video, âm thanh hoặc tệp 3D tùy biến thể Wan 2.2 bạn đã gọi.
Trường hợp sử dụng
Bạn có thể xây dựng gì với Wan 2.2
Các quy trình phổ biến mà nhà phát triển và người sáng tạo dùng API Wan 2.2 cho.
Wan 2.2 Animate — hoạt hình nhân vật tới 120 giây
wavespeed-ai/wan-2.2/animate là "mô hình hoạt hình & thay thế nhân vật hợp nhất, tái hiện chuyển động và biểu cảm; tạo video 720p dài tới 120 giây." Theo danh mục. Dài hơn đáng kể so với hầu hết công cụ hoạt hình điều khiển bằng tư thế.
Speech-to-Video tới 10 phút
wavespeed-ai/wan-2.2/speech-to-video "biến ảnh và giọng nói thành video độ trung thực cao với chuyển động khuôn mặt và cơ thể chân thực; hỗ trợ clip dài tới 10 phút ở 480p." Hữu ích cho nội dung nói chuyện dài.
Video Edit với thay đổi theo prompt
wavespeed-ai/wan-2.2/video-edit cho phép sửa video bằng prompt văn bản (ví dụ trong danh mục: đổi trang phục hoặc nhân vật). Hỗ trợ 480p và 720p, dài tới 120 giây.
Fun-Control với giấy phép Apache 2.0
wavespeed-ai/wan-2.2/fun-control dùng "Control Code và đầu vào đa phương thức để tạo video điều khiển theo cài sẵn dài tới 120 giây ở 720p; phát hành theo Apache 2.0 cho mục đích thương mại." Giấy phép Apache 2.0 là điểm khác biệt thực sự cho pipeline thương mại.
Huấn luyện LoRA (nhanh gấp 10 lần)
wavespeed-ai/wan-2.2-image-lora-trainer cho LoRA ảnh, wavespeed-ai/wan-2.2-i2v-lora-trainer cho LoRA I2V. Theo danh mục: "huấn luyện nhanh gấp 10 lần." Hỗ trợ huấn luyện phong cách, nhân vật, đối tượng, chuyển động, hành động và hiệu ứng video. Tải lên tệp ZIP để bắt đầu.
Image-to-Video ở nhiều cỡ (5B / A14B)
Chọn cỡ mô hình: 5B (nhỏ hơn) để ưu tiên tốc độ/chi phí, hoặc A14B cho chất lượng đầy đủ. i2v tiêu chuẩn cho 480p; các biến thể hỗ trợ LoRA; biến thể siêu nhanh.
Mẹo
Mẹo viết prompt cho Wan 2.2
Lời khuyên thực tế để có kết quả tốt hơn từ Wan 2.2 — rút ra từ các mẫu hiệu quả trên các mô hình video trong các pipeline sản xuất.
- 01
Chọn biến thể phù hợp tác vụ
Wan 2.2 có các endpoint chuyên biệt thay vì một mô hình đa dụng. Animate cho chuyển động điều khiển bằng tư thế, video-edit cho thay đổi có mục tiêu, speech-to-video cho nội dung nói chuyện, image-to-video cho tạo nội dung chung. Hãy khớp endpoint với tác vụ — đầu ra tốt hơn đáng kể so với bắt một mô hình đa dụng làm mọi việc.
- 02
Huấn luyện LoRA để nhất quán ở quy mô sản xuất
Các endpoint huấn luyện LoRA của Wan 2.2 là tính năng API hàng đầu, không phải công cụ phụ. Với các dự án cần danh tính lặp lại qua hàng trăm lần tạo (linh vật thương hiệu, nhân vật lặp lại, phong cách đặc trưng), hãy huấn luyện LoRA một lần rồi gọi endpoint suy luận LoRA cho mỗi lần tạo.
- 03
Trọng số mở cho phép tinh chỉnh
Các đối thủ trọng số đóng khóa bạn vào hành vi mô hình của nhà cung cấp. Nền tảng của Wan 2.2 là trọng số mở, nên khi gặp giới hạn mà mô hình gốc không giải quyết được, việc tinh chỉnh là một lựa chọn — hầu hết API video thương mại khác không cung cấp con đường này.
- 04
Kết hợp Animate với Kling Motion Control
Cả hai đều có hoạt hình điều khiển bằng tư thế với các đánh đổi khác nhau. Wan 2.2 Animate có hỗ trợ LoRA và trọng số mở; Kling Motion Control giữ danh tính tốt hơn. Hãy chọn tùy việc tinh chỉnh hay chất lượng danh tính quan trọng hơn.
Giá
Giá API Wan 2.2
Giá tính theo từng kết quả đầu ra. Khoản phí cuối cùng thay đổi theo các tham số bạn đặt trong playground của từng biến thể (độ phân giải, thời lượng, số kết quả, tham chiếu).
API
Gọi API Wan 2.2
Đăng ký API key tại wavespeed.ai/accesskey, rồi gửi prediction qua REST. Playground tạo sẵn mã mẫu có thể dán ngay cho mọi tổ hợp đầu vào.
POSThttps://api.wavespeed.ai/api/v3/wavespeed-ai/wan-2.2/image-to-video
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/wavespeed-ai/wan-2.2/image-to-video" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $WAVESPEED_API_KEY" \
-d '{
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"image": "https://interactive-examples.mdn.mozilla.net/media/cc0-images/painted-hand-298-332.jpg",
"resolution": "480p",
"duration": 5
}')
TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
-H "Authorization: Bearer $WAVESPEED_API_KEY")
RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
case "$STATUS" in
completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
failed|cancelled|timeout|deleted) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
*) sleep 2 ;;
esac
doneconst submitUrl = "https://api.wavespeed.ai/api/v3/wavespeed-ai/wan-2.2/image-to-video";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');
async function requestJson(url, options = {}) {
const response = await fetch(url, options);
if (!response.ok) throw new Error(await response.text());
return response.json();
}
// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
method: "POST",
headers: {
"Authorization": `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"image": "https://interactive-examples.mdn.mozilla.net/media/cc0-images/painted-hand-298-332.jpg",
"resolution": "480p",
"duration": 5
}),
});
const task = body.data ?? body;
const resultUrl = `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;
// 2. Poll until the prediction finishes.
while (true) {
const resultBody = await requestJson(resultUrl, {
headers: { "Authorization": `Bearer ${apiKey}` },
});
const result = resultBody.data ?? resultBody;
if (result.status === "completed") {
console.log(result.outputs);
break;
}
if (["failed", "cancelled", "timeout", "deleted"].includes(result.status)) throw new Error(JSON.stringify(result));
await new Promise(resolve => setTimeout(resolve, 2000));
}import json
import os
import time
from urllib.request import Request, urlopen
api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"image": "https://interactive-examples.mdn.mozilla.net/media/cc0-images/painted-hand-298-332.jpg",
"resolution": "480p",
"duration": 5
}
def request_json(url, data=None):
request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
with urlopen(request) as response:
return json.load(response)
# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/wavespeed-ai/wan-2.2/image-to-video", json.dumps(payload).encode())
task = body.get("data", body)
result_url = f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"
# 2. Poll until the prediction finishes.
while True:
result_body = request_json(result_url)
result = result_body.get("data", result_body)
status = result.get("status")
if status == "completed":
print(result.get("outputs", []))
break
if status in {"failed", "cancelled", "timeout", "deleted"}:
raise RuntimeError(result)
time.sleep(2)So sánh
Wan 2.2 so với các lựa chọn khác
Khi nào nên chọn Wan 2.2 thay vì các mô hình tương tự trên WaveSpeedAI.
Wan 2.2 so với Wan 2.7
Wan 2.7 (alibaba/wan-2.7/*) là kiến trúc mới hơn của Alibaba, gộp reference-to-video, video-edit, image-edit và text-to-image trong một dòng — bộ công cụ đa phương thức rộng hơn. Wan 2.2 (các biến thể WaveSpeedAI) có các endpoint chuyên biệt — Animate (120 giây), Speech-to-Video (10 phút), Fun-Control (Apache 2.0), trình huấn luyện LoRA — mà 2.7 không cung cấp.
Wan 2.2 so với Seedance 2.0
Seedance 2.0 cho đầu ra chuẩn Hollywood với âm thanh gốc ở mọi biến thể. Wan 2.2 thắng ở độ đa dạng biến thể (hơn 35 endpoint) và khả năng huấn luyện LoRA — lựa chọn đúng khi bạn cần một năng lực chuyên biệt (Animate, Speech-to-Video, Fun-Control) thay vì một mô hình video đa dụng.
Wan 2.2 so với Kling 3.0 Motion Control
Cả hai đều có hoạt hình nhân vật điều khiển bằng tư thế. Wan 2.2 Animate tạo clip 720p dài 120 giây và hỗ trợ tinh chỉnh LoRA. Kling Motion Control bị giới hạn theo độ dài video tham chiếu (3-30 giây) nhưng được huấn luyện trên kho video của Kuaishou nên có tiên nghiệm chuyển động mạnh hơn.
Câu hỏi thường gặp
Wan 2.2 API — Câu hỏi thường gặp
Giá, giấy phép, tích hợp — những câu hỏi phổ biến về việc chạy Wan 2.2 trên WaveSpeedAI.
API Wan 2.2 là gì?
Wan 2.2 là mô hình tạo video của Alibaba, được cung cấp dưới dạng REST API trên WaveSpeedAI. Wan 2.2 của Alibaba — bộ công cụ video mã nguồn mở trọng số triển khai trên WaveSpeedAI với hơn 35 biến thể chính chủ: Animate (hoạt hình nhân vật 120 giây), Video Edit, Speech-to-Video (10 phút theo âm thanh), Fun-Control (giấy phép Apache 2.0), cùng image-to-video và text-to-video ở nhiều cỡ mô hình (5B, A14B) và độ phân giải (480p / 720p). Bạn có thể gọi bằng lập trình hoặc thử từ playground ở liên kết phía trên.
Làm sao để gọi API Wan 2.2?
Đăng ký tài khoản WaveSpeedAI, sao chép API key từ /accesskey, rồi POST tới https://api.wavespeed.ai/api/v3/wavespeed-ai/wan-2.2/image-to-video với đầu vào dưới dạng JSON. Endpoint trả về prediction id. Hãy poll endpoint kết quả khoảng mỗi 2 giây, tăng khoảng cách với các tác vụ chạy lâu, và dừng ở bất kỳ trạng thái kết thúc nào. Ví dụ Python / Node.js / cURL hướng tới môi trường production ở phía trên.
API Wan 2.2 có giá bao nhiêu?
Wan 2.2 bắt đầu từ $0.02 mỗi lượt chạy. Chi phí chính xác thay đổi theo các tham số bạn đặt (độ phân giải, thời lượng, số kết quả, tham chiếu). Phần xem trước chi phí trực tiếp cạnh nút Tạo trong playground hiển thị giá chính xác cho đầu vào hiện tại của bạn.
Có những biến thể Wan 2.2 nào?
WaveSpeedAI cung cấp 32 endpoint Wan 2.2 đang hoạt động: wavespeed-ai/wan-2.2/image-to-video-lora, wavespeed-ai/wan-2.2/image-to-video, wavespeed-ai/wan-2.2/animate, wavespeed-ai/wan-2.2/video-edit, wavespeed-ai/wan-2.2/image-to-image, wavespeed-ai/wan-2.2/speech-to-video, wavespeed-ai/wan-2.2/text-to-image-lora, wavespeed-ai/wan-2.2/fun-control, và nhiều hơn nữa. Mỗi biến thể có trang playground và giá riêng.
Tôi có thể dùng kết quả của Wan 2.2 cho mục đích thương mại không?
Quyền sử dụng thương mại tuân theo giấy phép mô hình của Alibaba. Hầu hết mô hình của Alibaba cho phép dùng kết quả đầu ra cho mục đích thương mại; hãy xem trang playground của từng mô hình để biết tóm tắt giấy phép cụ thể, và Điều khoản dịch vụ của WaveSpeedAI cho các điều kiện cấp nền tảng.
Tại sao dùng Wan 2.2 trên WaveSpeedAI thay vì gọi trực tiếp?
Một API key + một tài khoản thanh toán cho Wan 2.2 VÀ hơn 1.000 mô hình AI khác từ các nhà cung cấp khác. Không cần thiết lập SDK cho từng hãng, không có hạn mức tốc độ riêng, không phải viết lại mã tích hợp cho từng hãng. Giá thường ngang bằng hoặc thấp hơn API trực tiếp của Alibaba.
Nhà cung cấp
Giới thiệu về Alibaba
Đội ngũ đứng sau Wan 2.2 và dòng mô hình Alibaba rộng hơn trên WaveSpeedAI.
Tongyi Lab của Alibaba tạo ra dòng mô hình video Wan và dòng LLM Qwen. Wan nổi bật nhờ được phát hành với trọng số mở, phủ nhiều biến thể (text-to-video, image-to-video, reference-to-video, video-edit, video-extend, image-edit, text-to-image) và luôn mạnh về độ ổn định chuyển động cũng như khả năng bám sát prompt với các prompt đa ngôn ngữ.
Bắt đầu xây dựng với Wan 2.2 trên WaveSpeedAI
Tặng credit dùng thử miễn phí khi đăng ký. Một API key cho hơn 1.000 mô hình AI từ Alibaba và mọi nhà cung cấp khác.