Vidu Q4 API
Shengshu Vidu Q4 — image-to-video and reference-to-video with 3-16 second clips, 540p to 4K output, and optional generated audio. Reference-to-video takes up to 12 reference images and up to 3 reference audio clips in one request.
Two endpoints: vidu/q4-preview/image-to-video animates a single image, and vidu/q4-preview/reference-to-video builds a scene from a prompt plus up to 12 reference images and 3 reference audio clips. Both take any whole-second duration from 3 to 16 seconds, five resolutions from 540p to 4K, and an audio toggle.
Overview
About the Vidu Q4 API
What Vidu Q4 does, how it fits in the Shengshu model lineup, and why teams reach for it.
Vidu Q4 is a video generation model from Shengshu, available through the WaveSpeedAI REST API. Shengshu Vidu Q4 — image-to-video and reference-to-video with 3-16 second clips, 540p to 4K output, and optional generated audio. Reference-to-video takes up to 12 reference images and up to 3 reference audio clips in one request.
Two endpoints: vidu/q4-preview/image-to-video animates a single image, and vidu/q4-preview/reference-to-video builds a scene from a prompt plus up to 12 reference images and 3 reference audio clips. Both take any whole-second duration from 3 to 16 seconds, five resolutions from 540p to 4K, and an audio toggle.
The Vidu Q4 family on WaveSpeedAI ships 2 REST endpoints covering Reference-To-Video, Image-To-Video workflows. Each variant carries its own pricing, parameter knobs, and example outputs — pick the one that matches your input modality and production constraints, or call several from the same API key to compose multi-step pipelines.
Run Vidu Q4 through the same API key, billing account, and rate-limit envelope you use for the other 1,000+ AI models on WaveSpeedAI. No separate vendor setup, no per-provider SDKs, no per-vendor rate-limit envelopes — one integration covers everything from text-to-image and text-to-video through audio synthesis, 3D generation, upscaling, and editing.
Specs
Vidu Q4 API capabilities and release status
The model-specific details developers search for before choosing an API: availability, expected output length, reference support, and the current live fallback.
Endpoints
2 variants
Image-to-video and reference-to-video.
Duration
3-16 s
Any whole number of seconds; 5 s is the default.
Resolution
540p to 4K
540p, 720p, 1080p, 2K, 4K; 720p is the default.
References
12 images + 3 audio
On reference-to-video, alongside a text prompt.
Endpoints
All Vidu Q4 API endpoints
2 Vidu Q4 endpoints available now on WaveSpeedAI — pick the variant that matches your workflow.
/filters:quality(82)/media/images/1790656175314160249_JiSjQ9AW.webp)
Q4 Preview Reference To Video
Vidu Q4 Reference-to-Video generates subject-consistent AI videos from text prompts, 1-12 reference images, and up to 3 reference audio clips, supporting character consistency, product videos, social content, brand assets, and reference-guided storytelling workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
/filters:quality(82)/media/images/1790656159881988529_32y4wP0a.webp)
Q4 Preview Image To Video
Vidu Q4 Image-to-Video animates a single reference image into high-quality AI video with prompt-guided motion, optional audio, 3-16 second duration, and output resolutions from 540P to 4K for product animation, social content, marketing creatives, cinematic visuals, and production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Examples
See Vidu Q4 in action
Real outputs generated by the Vidu Q4 API. Hover any video to preview, click to open the full-size viewer.
How to
How to use the Vidu Q4 API
Four steps from signup to a finished generation. Full Python, Node.js, and cURL examples are in the API section below.
- 01
Get an API key
Sign up for a WaveSpeedAI account and copy your API key from the dashboard. New accounts come with free starter credits — enough to run the playground a few dozen times before billing kicks in.
- 02
Submit a prediction
POST your input as JSON to https://api.wavespeed.ai/api/v3/vidu/q4-preview/image-to-video. The endpoint returns a prediction id immediately — generations are async so you don't hold an open connection during inference.
- 03
Poll for completion
GET https://api.wavespeed.ai/api/v3/predictions/{request_id}/result. Return outputs on completed; stop with an error on failed, cancelled, timeout, or deleted; continue polling for every other status.
- 04
Read the output URL
Once status is "completed", read the URL from data.outputs[0]. The URL points to your generated media on the WaveSpeedAI CDN — image, video, audio, or 3D file depending on the Vidu Q4 variant you called.
Use cases
What you can build with Vidu Q4
Common workflows developers and creators use the Vidu Q4 API for.
Product animation from one photo
vidu/q4-preview/image-to-video turns a single product shot into a promotional clip. Describe the camera move and the motion you want; the image stays the visual anchor.
Characters that stay on model
vidu/q4-preview/reference-to-video accepts up to 12 reference images, so a recurring character, mascot, or spokesperson can be shown from several angles and stay recognizable across clips.
Scenes guided by sound
Add up to 3 reference audio clips through the audios field alongside your images, and keep generate_audio on to get a clip with sound in a single call.
Delivery-ready 4K
Both endpoints output 540p, 720p, 1080p, 2K, or 4K. Draft at 540p or 720p, then re-run the approved prompt at the delivery resolution without a separate upscaling pass.
Clips sized to the beat
Duration is any whole number of seconds from 3 to 16, so a 3-second loop, a 6-second bumper, and a 15-second spot all come from the same endpoint.
Tips
Tips for prompting Vidu Q4
Practical advice for getting better outputs from Vidu Q4 — drawn from the patterns that work across video models in production pipelines.
- 01
Use a clean, well-lit source image
For vidu/q4-preview/image-to-video the image sets the subject, framing, and style. Pick a sharp image with one clear subject, then use the prompt for motion and camera direction instead of re-describing what is already visible.
- 02
Give each reference a job
With vidu/q4-preview/reference-to-video, use references that belong in the scene — the character, the product, the setting — and say in the prompt how each one should appear rather than sending loosely related images.
- 03
Match the action to the duration
Ask for an amount of action that fits the clip length. A 3-5 second clip suits one move or gesture; save multi-step action for 10-16 seconds.
- 04
Draft low, deliver high
Explore ideas at 540p or 720p with short durations, then re-run the chosen prompt at 1080p, 2K, or 4K for delivery.
- 05
Decide on audio and prompt rewriting up front
generate_audio is on by default; turn it off when you will add your own soundtrack. On image-to-video, set enable_prompt_expansion to false when you want your prompt used as written.
Pricing
Vidu Q4 API pricing
Pricing is per-output. The final charge scales with the parameters you set in each variant's playground (resolution, duration, output count, references).
| Endpoint | Type | Starting price |
|---|---|---|
| vidu/ | reference-to-video | $0.25 |
| vidu/ | image-to-video | $0.25 |
API
Call the Vidu Q4 API
Sign up for an API key at wavespeed.ai/accesskey, then submit a prediction via REST. The playground generates ready-to-paste samples for any combination of inputs.
POSThttps://api.wavespeed.ai/api/v3/vidu/q4-preview/image-to-video
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/vidu/q4-preview/image-to-video" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $WAVESPEED_API_KEY" \
-d '{
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"image": "https://interactive-examples.mdn.mozilla.net/media/cc0-images/painted-hand-298-332.jpg",
"resolution": "720p",
"duration": 5,
"enable_prompt_expansion": true,
"generate_audio": true
}')
TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
-H "Authorization: Bearer $WAVESPEED_API_KEY")
RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
case "$STATUS" in
completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
failed|cancelled|timeout|deleted) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
*) sleep 2 ;;
esac
doneconst submitUrl = "https://api.wavespeed.ai/api/v3/vidu/q4-preview/image-to-video";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');
async function requestJson(url, options = {}) {
const response = await fetch(url, options);
if (!response.ok) throw new Error(await response.text());
return response.json();
}
// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
method: "POST",
headers: {
"Authorization": `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"image": "https://interactive-examples.mdn.mozilla.net/media/cc0-images/painted-hand-298-332.jpg",
"resolution": "720p",
"duration": 5,
"enable_prompt_expansion": true,
"generate_audio": true
}),
});
const task = body.data ?? body;
const resultUrl = `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;
// 2. Poll until the prediction finishes.
while (true) {
const resultBody = await requestJson(resultUrl, {
headers: { "Authorization": `Bearer ${apiKey}` },
});
const result = resultBody.data ?? resultBody;
if (result.status === "completed") {
console.log(result.outputs);
break;
}
if (["failed", "cancelled", "timeout", "deleted"].includes(result.status)) throw new Error(JSON.stringify(result));
await new Promise(resolve => setTimeout(resolve, 2000));
}import json
import os
import time
from urllib.request import Request, urlopen
api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"image": "https://interactive-examples.mdn.mozilla.net/media/cc0-images/painted-hand-298-332.jpg",
"resolution": "720p",
"duration": 5,
"enable_prompt_expansion": True,
"generate_audio": True
}
def request_json(url, data=None):
request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
with urlopen(request) as response:
return json.load(response)
# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/vidu/q4-preview/image-to-video", json.dumps(payload).encode())
task = body.get("data", body)
result_url = f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"
# 2. Poll until the prediction finishes.
while True:
result_body = request_json(result_url)
result = result_body.get("data", result_body)
status = result.get("status")
if status == "completed":
print(result.get("outputs", []))
break
if status in {"failed", "cancelled", "timeout", "deleted"}:
raise RuntimeError(result)
time.sleep(2)Compare
Vidu Q4 vs alternatives
When to pick Vidu Q4 over similar models on WaveSpeedAI.
Vidu Q4 vs Vidu Q3
Vidu Q3 reference-to-video takes 1-4 reference images; Q4 takes up to 12 images plus up to 3 audio references, and both Q4 endpoints reach 4K. Q3 remains the pick for text-to-video, start-end keyframe interpolation, and its Pro and Turbo tiers.
Vidu Q4 vs Seedance 2.5
Seedance 2.5 goes up to 30-second single shots and adds video editing and extension endpoints. Vidu Q4 focuses on image- and reference-driven generation with output up to 4K.
Vidu Q4 vs Wan 3.0
Wan 3.0 offers text-, image-, and reference-to-video up to 30 seconds at up to 1080p. Vidu Q4 tops out at 16 seconds but adds 2K and 4K output and up to 12 reference images.
FAQ
Vidu Q4 API — Frequently asked questions
Pricing, license, integration — common questions about running Vidu Q4 on WaveSpeedAI.
What is the Vidu Q4 API?
Vidu Q4 is a Shengshu video generation model exposed as a REST API on WaveSpeedAI. Shengshu Vidu Q4 — image-to-video and reference-to-video with 3-16 second clips, 540p to 4K output, and optional generated audio. Reference-to-video takes up to 12 reference images and up to 3 reference audio clips in one request. You can call it programmatically or try it from the playground linked above.
How do I call the Vidu Q4 API?
Sign up for a WaveSpeedAI account, copy your API key from /accesskey, then POST to https://api.wavespeed.ai/api/v3/vidu/q4-preview/image-to-video with your input as JSON. The endpoint returns a prediction id. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. Production-oriented Python / Node.js / cURL examples are above.
How much does the Vidu Q4 API cost?
Vidu Q4 starts at $0.25 per run. The exact cost scales with the parameters you set (resolution, duration, output count, references). The live cost preview next to the Generate button in the playground shows the exact price for your current input.
Which Vidu Q4 variants are available?
WaveSpeedAI hosts 2 live Vidu Q4 endpoints: vidu/q4-preview/reference-to-video, vidu/q4-preview/image-to-video. Each variant has its own playground page and pricing.
Can I use Vidu Q4 outputs commercially?
Commercial usage rights follow the Shengshu model license. Most Shengshu models permit commercial output use; see each model's playground page for the specific license summary, and WaveSpeedAI's Terms of Service for platform-level conditions.
Why use Vidu Q4 on WaveSpeedAI instead of going direct?
One API key + one billing account across Vidu Q4 AND 1,000+ other AI models from other providers. No per-vendor SDK setup, no separate rate-limit envelopes, no rewrite-per-vendor integration code. Pricing is typically at parity with or below Shengshu's direct API.
Provider
About Shengshu
The team behind Vidu Q4 and the broader Shengshu model lineup on WaveSpeedAI.
Shengshu Technology is a Chinese AI lab spun out of Tsinghua University, behind the Vidu family of video generation models. Vidu Q4, the newest generation, adds image-to-video and reference-to-video with up to 12 reference images, up to 3 reference audio clips, and output up to 4K. Vidu Q3 ships text-to-video, image-to-video, reference-to-video (1-4 reference images for multi-entity consistency), and start-end-to-video (keyframe interpolation between two stills) across Standard, Pro, and Turbo tiers. Some variants support up to 16-second outputs.
Start building with Vidu Q4 on WaveSpeedAI
Free starter credits on signup. One API key across 1,000+ AI models from Shengshu and every other provider.