Wan 3.0 API
Alibaba Wan 3.0 — cinematic video in a single generation pass of up to 30 seconds, double the ceiling of Wan 2.7. Text-to-video, image-to-video with first and last frame guidance, and reference-to-video that takes image, video, and audio references together. 480p / 720p / 1080p output, optional generated audio, and a thinking mode for higher-quality motion and scene continuity.
Three endpoints on one family: text-to-video from a prompt, image-to-video that animates a first frame with optional last-frame guidance, and reference-to-video driven by multimodal references. Every variant takes a 2–30 second duration (5s default), 480p / 720p / 1080p output, five aspect ratios (16:9, 9:16, 1:1, 4:3, 3:4), an optional audio toggle, and deep-thinking controls.
About the Wan 3.0 API
What Wan 3.0 does, how it fits in the Alibaba model lineup, and why teams reach for it.
Wan 3.0 is a video generation model from Alibaba, available through the WaveSpeedAI REST API. Alibaba Wan 3.0 — cinematic video in a single generation pass of up to 30 seconds, double the ceiling of Wan 2.7. Text-to-video, image-to-video with first and last frame guidance, and reference-to-video that takes image, video, and audio references together. 480p / 720p / 1080p output, optional generated audio, and a thinking mode for higher-quality motion and scene continuity.
Three endpoints on one family: text-to-video from a prompt, image-to-video that animates a first frame with optional last-frame guidance, and reference-to-video driven by multimodal references. Every variant takes a 2–30 second duration (5s default), 480p / 720p / 1080p output, five aspect ratios (16:9, 9:16, 1:1, 4:3, 3:4), an optional audio toggle, and deep-thinking controls.
The Wan 3.0 family on WaveSpeedAI ships 3 REST endpoints covering Image-To-Video, Text-To-Video workflows. Each variant carries its own pricing, parameter knobs, and example outputs — pick the one that matches your input modality and production constraints, or call several from the same API key to compose multi-step pipelines.
Run Wan 3.0 through the same API key, billing account, and rate-limit envelope you use for the other 1,000+ AI models on WaveSpeedAI. No separate vendor setup, no per-provider SDKs, no per-vendor rate-limit envelopes — one integration covers everything from text-to-image and text-to-video through audio synthesis, 3D generation, upscaling, and editing.
Wan 3.0 API capabilities and release status
The model-specific details developers search for before choosing an API: availability, expected output length, reference support, and the current live fallback.
Max duration
30 seconds
Every variant accepts 2–30 seconds in one generation pass, 5 seconds by default.
Resolution
480p / 720p / 1080p
Same three tiers across text-to-video, image-to-video, and reference-to-video.
Endpoints
3 variants
text-to-video, image-to-video (first + last frame), reference-to-video (image, video, audio references).
Aspect ratios
5 options
16:9, 9:16, 1:1, 4:3, and 3:4 on every variant.
All Wan 3.0 API endpoints
3 Wan 3.0 endpoints available now on WaveSpeedAI — pick the variant that matches your workflow.

Wan 3.0 Reference To Video
Wan 3.0 Reference to Video creates coherent videos from prompts and multimodal references, including images, videos, and audio, with flexible 2-30 second duration and aspect ratio control for subject consistency, motion guidance, timing control, and scene continuity. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Wan 3.0 Image To Video
Wan 3.0 Image to Video animates a first-frame image into a cinematic video, with optional last-frame guidance, flexible 2-30 second duration, aspect ratio control, optional audio, and deep-thinking controls for high-quality motion and scene continuity. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Wan 3.0 Text To Video
Wan 3.0 Text to Video generates cinematic videos from text prompts, with flexible 2-30 second duration, aspect ratio control, optional audio, and deep-thinking controls for high-quality video generation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
See Wan 3.0 in action
Real outputs generated by the Wan 3.0 API. Hover any video to preview, click to open the full-size viewer.
How to use the Wan 3.0 API
Four steps from signup to a finished generation. Full Python, Node.js, and cURL examples are in the API section below.
- 1
Get an API key
Sign up for a WaveSpeedAI account and copy your API key from the dashboard. New accounts come with free starter credits — enough to run the playground a few dozen times before billing kicks in.
- 2
Submit a prediction
POST your input as JSON to https://api.wavespeed.ai/api/v3/alibaba/wan-3.0/text-to-video. The endpoint returns a prediction id immediately — generations are async so you don't hold an open connection during inference.
- 3
Poll for completion
GET https://api.wavespeed.ai/api/v3/predictions/{request_id}/result. Start around every 2 seconds, then increase toward 5-10 seconds for long-running tasks to reduce unnecessary requests. Stop on completed, failed, cancelled, or timeout.
- 4
Read the output URL
Once status is"completed", read the URL from data.outputs[0]. The URL points to your generated media on the WaveSpeedAI CDN — image, video, audio, or 3D file depending on the Wan 3.0 variant you called.
What you can build with Wan 3.0
Common workflows developers and creators use the Wan 3.0 API for.
Single-pass shots up to 30 seconds
Every Wan 3.0 variant accepts a 2–30 second duration. Because the clip comes out of one generation pass, a continuous camera move or one-take shot language holds across the whole run instead of being stitched from shorter segments. Wan 2.7 tops out at 15 seconds.
Text-to-video with deep-thinking controls
alibaba/wan-3.0/text-to-video turns a prompt into a cinematic clip. Catalog framing: aspect ratio control, optional audio, and deep-thinking controls for high-quality video generation. Pick this when nothing is locked yet and the shot is being built from description alone.
Image-to-video with first and last frame
alibaba/wan-3.0/image-to-video animates a first-frame image, with optional last-frame guidance via the last_image field. Useful when both the opening and closing composition are already approved and the generation only needs to fill the motion between them.
Reference-to-video with image, video, and audio refs
alibaba/wan-3.0/reference-to-video accepts reference_images, reference_videos, and reference_audios in the same request — for subject consistency, motion guidance, timing control, and scene continuity. The audio reference channel is what separates it from the image-only reference variants elsewhere in the Wan family.
Draft at 480p, deliver at 1080p
All three variants expose 480p, 720p, and 1080p. Wan 2.7 starts at 720p, so 3.0 adds a genuinely cheap draft tier — iterate prompt and timing at 480p, then re-run the settled prompt at 1080p for delivery.
Vertical, square, and broadcast framing
Five aspect ratios (16:9, 9:16, 1:1, 4:3, 3:4) are available on every variant, so the same prompt can be re-run for a landscape hero cut, a 9:16 social edit, and a 1:1 feed placement without cropping in post.
Tips for prompting Wan 3.0
Practical advice for getting better outputs from Wan 3.0 — drawn from the patterns that work across video models in production pipelines.
Be specific about camera moves
Mention concrete cinematography vocabulary — orbit, dolly-in, push-in, pan-left, crane shot, handheld follow. Generic prompts produce static or arbitrary camera choices; named camera moves map directly to motion intent in the model's training data and dramatically improve shot quality.
Anchor character identity with reference images
If your prompt depends on a specific person, character, or product, upload a reference image alongside the prompt. Without a reference, identity drifts across frames and across shots — the same character ends up looking like a slightly different person each generation.
Describe lighting and time of day
Lighting cues like 'golden hour, soft warm directional light' or 'overcast diffused light, slate-grey sky' improve quality and consistency far more than vague quality modifiers. Lighting is one of the strongest priors the model conditions on.
Use negative prompts to suppress common failure modes
Useful negatives for video: 'frame flicker, motion blur, watermark, text artifacts, distorted hands, low resolution, jpeg compression'. Negative prompts cost nothing and noticeably reduce the rate of generations you'd otherwise re-roll.
Pick the shortest duration that captures your beat
Most prompts work best at 5-8 seconds. Longer clips amplify temporal inconsistencies (subject morphing, environment drift). If you need a 20-second sequence, generate three 6-8 second clips and edit them together — quality stays higher than one long generation.
Match aspect ratio to platform up front
9:16 for TikTok / Reels / Shorts, 16:9 for landscape feeds and YouTube, 1:1 for post grids. Models train slightly differently per aspect ratio — cropping a 16:9 to 9:16 after the fact loses both fidelity and the composition the model intended.
Wan 3.0 API pricing
Pricing is per-output. The final charge scales with the parameters you set in each variant's playground (resolution, duration, output count, references).
| Endpoint | Type | Starting price |
|---|---|---|
| alibaba/wan-3.0/reference-to-video | image-to-video | $0.60 |
| alibaba/wan-3.0/image-to-video | image-to-video | $0.60 |
| alibaba/wan-3.0/text-to-video | text-to-video | $0.60 |
Call the Wan 3.0 API
Sign up for an API key at wavespeed.ai/accesskey, then submit a prediction via REST. The playground generates ready-to-paste samples for any combination of inputs.
HTTP example
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/alibaba/wan-3.0/text-to-video" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $WAVESPEED_API_KEY" \
-d '{}')
TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL=$(printf '%s' "$TASK" | jq -r '.urls.get // empty')
if [ -z "$RESULT_URL" ]; then
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
fi
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
-H "Authorization: Bearer $WAVESPEED_API_KEY")
RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
case "$STATUS" in
completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
failed|cancelled|timeout) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
created|processing) sleep 2 ;;
*) printf 'Unexpected status: %s
' "$STATUS" >&2; exit 1 ;;
esac
doneNode.js example
const submitUrl = "https://api.wavespeed.ai/api/v3/alibaba/wan-3.0/text-to-video";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');
async function requestJson(url, options = {}) {
const response = await fetch(url, options);
if (!response.ok) throw new Error(await response.text());
return response.json();
}
// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
method: "POST",
headers: {
"Authorization": `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({}),
});
const task = body.data ?? body;
const resultUrl = task.urls?.get ||
`https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;
// 2. Poll until the prediction finishes.
while (true) {
const resultBody = await requestJson(resultUrl, {
headers: { "Authorization": `Bearer ${apiKey}` },
});
const result = resultBody.data ?? resultBody;
if (result.status === "completed") {
console.log(result.outputs);
break;
}
if (["failed", "cancelled", "timeout"].includes(result.status)) throw new Error(JSON.stringify(result));
if (!["created", "processing"].includes(result.status)) throw new Error("Unexpected status: " + result.status);
await new Promise(resolve => setTimeout(resolve, 2000));
}Python example
import json
import os
import time
from urllib.request import Request, urlopen
api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {}
def request_json(url, data=None):
request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
with urlopen(request) as response:
return json.load(response)
# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/alibaba/wan-3.0/text-to-video", json.dumps(payload).encode())
task = body.get("data", body)
result_url = task.get("urls", {}).get("get") or f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"
# 2. Poll until the prediction finishes.
while True:
result_body = request_json(result_url)
result = result_body.get("data", result_body)
status = result.get("status")
if status == "completed":
print(result.get("outputs", []))
break
if status in {"failed", "cancelled", "timeout"}:
raise RuntimeError(result)
if status not in {"created", "processing"}:
raise RuntimeError(f"Unexpected status: {status}")
time.sleep(2)Wan 3.0 vs alternatives
When to pick Wan 3.0 over similar models on WaveSpeedAI.
Wan 3.0 vs Wan 2.7
Wan 3.0 doubles the duration ceiling — 30 seconds against 15 — and adds a 480p tier plus explicit enable_audio and thinking_mode controls that 2.7 does not expose. Wan 2.7 is still the broader family: it ships video-edit, video-extend, image-edit, and text-to-image variants that 3.0's three endpoints do not cover.
Wan 3.0 vs Seedance 2.0
Seedance 2.0 reaches 4K and Wan 3.0 stops at 1080p, but Seedance caps at 15 seconds. The trade is resolution against shot length: pick Seedance for a short cut that has to hold up at 4K, Wan 3.0 for a long continuous take.
Wan 3.0 vs Happy Horse 1.1
Happy Horse 1.1, also from Alibaba, runs 720p/1080p at up to 15 seconds and is positioned around smooth camera movement. Wan 3.0 wins on duration and adds the 480p draft tier; Happy Horse covers video-extend, which Wan 3.0 does not ship.
Wan 3.0 API — Frequently asked questions
Pricing, license, integration — common questions about running Wan 3.0 on WaveSpeedAI.
How long can a Wan 3.0 video be?
Between 2 and 30 seconds, set per request, with 5 seconds as the default. The full clip is produced in a single generation pass rather than stitched together, so continuous camera movement holds across the whole duration.
What is the difference between Wan 3.0 and Wan 2.7?
Wan 3.0 doubles the maximum duration from 15 to 30 seconds, adds a 480p output tier below 720p, and exposes explicit enable_audio and thinking_mode controls. Wan 2.7 covers more task types — it adds video-edit, video-extend, image-edit, and text-to-image endpoints that Wan 3.0 does not.
Can Wan 3.0 generate audio with the video?
Yes. Every variant exposes an optional audio toggle. The reference-to-video variant additionally accepts audio references alongside image and video references, which can be used for timing control and scene continuity.
Which Wan 3.0 endpoint should I use?
Use text-to-video when the shot is being built from a description alone, image-to-video when the opening frame is already approved (optionally supplying a last frame to fix the ending composition), and reference-to-video when identity, motion, or timing should be guided by existing images, videos, or audio.
What is the Wan 3.0 API?
Wan 3.0 is a Alibaba video generation model exposed as a REST API on WaveSpeedAI. Alibaba Wan 3.0 — cinematic video in a single generation pass of up to 30 seconds, double the ceiling of Wan 2.7. Text-to-video, image-to-video with first and last frame guidance, and reference-to-video that takes image, video, and audio references together. 480p / 720p / 1080p output, optional generated audio, and a thinking mode for higher-quality motion and scene continuity. You can call it programmatically or try it from the playground linked above.
How do I call the Wan 3.0 API?
Sign up for a WaveSpeedAI account, copy your API key from /accesskey, then POST to https://api.wavespeed.ai/api/v3/alibaba/wan-3.0/text-to-video with your input as JSON. The endpoint returns a prediction id. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. Production-oriented Python / Node.js / cURL examples are above.
How much does the Wan 3.0 API cost?
Wan 3.0 starts at $0.60 per run. The exact cost scales with the parameters you set (resolution, duration, output count, references). The live cost preview next to the Generate button in the playground shows the exact price for your current input.
Which Wan 3.0 variants are available?
WaveSpeedAI hosts 3 live Wan 3.0 endpoints: alibaba/wan-3.0/reference-to-video, alibaba/wan-3.0/image-to-video, alibaba/wan-3.0/text-to-video. Each variant has its own playground page and pricing.
Can I use Wan 3.0 outputs commercially?
Commercial usage rights follow the Alibaba model license. Most Alibaba models permit commercial output use; see each model's playground page for the specific license summary, and WaveSpeedAI's Terms of Service for platform-level conditions.
Why use Wan 3.0 on WaveSpeedAI instead of going direct?
One API key + one billing account across Wan 3.0 AND 1,000+ other AI models from other providers. No per-vendor SDK setup, no separate rate-limit envelopes, no rewrite-per-vendor integration code. Pricing is typically at parity with or below Alibaba's direct API.
About Alibaba
The team behind Wan 3.0 and the broader Alibaba model lineup on WaveSpeedAI.
Alibaba's Tongyi Lab produces the Wan family of video models and the Qwen family of LLMs. Wan is notable for being released with open weights, broad variant coverage (text-to-video, image-to-video, reference-to-video, video-edit, video-extend, image-edit, text-to-image), and consistent strength on motion stability and prompt adherence across multilingual prompts.
Start building with Wan 3.0 on WaveSpeedAI
Free starter credits on signup. One API key across 1,000+ AI models from Alibaba and every other provider.