Vidu Text to Video converts text prompts into high-quality 720p videos with exceptional visual fidelity and diverse motion dynamics. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Idle
$0.2per run·~50 / $10
In a narrow, rain-slicked alley of a futuristic Tokyo, steam rises from grates. A cyborg samurai, wearing a modern tactical kimono, slowly draws his high-frequency energy katana from its sheath. The blade hums and glows with intense pink energy, casting vibrant reflections on the wet ground and his chrome prosthetic arm. Extreme detail, cinematic atmosphere.
Extreme macro shot of a miniature Victorian city inside a glass bottle. A tiny steam train is chugging along a track built into the bottle's inner wall, puffing out cotton-like steam. Sunlight from outside shines through the glass, creating warm light spots on the tiny streets.
In a futuristic cyberpunk city, neon billboards flicker in the rain. A sleek, high-speed hovercar weaves between skyscrapers, dodging laser fire. Rain streaks across its windows, with other flying vehicles and blurred city lights in the background. Cinematic, high-speed motion shot.
A little girl in a futuristic holographic zoo curiously reaches out, trying to touch a lifelike, giant holographic whale. The whale swims gracefully through the air, emitting a soft blue glow. Other holographic creatures and excited visitors are visible in the background, all in awe. Soft, dreamy camera shot.
In a cozy, retro cat cafe, a ginger cat wearing a tiny apron clumsily pushes a tray with coffee and cookies. It waddles past several sunlit tables, as amused customers take out their phones to snap pictures. Camera follows the cat.
Focus on a windowsill in a city apartment, where a tiny seed begins to sprout and grow rapidly. Vines and green leaves spread outwards from the window, gradually covering the surrounding grey concrete walls, forming a miniature urban forest. Time-lapse effect, with alternating sunlight and rain.
In a dimly lit, smoky, vintage jazz bar, a female singer in a shimmering crimson gown sings soulfully with her eyes closed, holding a classic microphone. The camera slowly pushes in on her profile, with the soft blur of the band and the warm gleam of a saxophone in the background.
At dusk, with the sky painted in orange and purple hues, a couple dances a passionate tango barefoot on the wet, empty sand. The waves gently lap at their feet as their bodies create elegant silhouettes against the setting sun. The camera circles them.
Inside a speeding, vintage luxury train car, a man sits on a velvet seat, gazing at the passing landscape. His contemplative face is reflected in the window pane. Light from outside periodically sweeps across him and the glass of whiskey in his hand.
Vidu Text-to-Video transforms your text prompts into high-quality, cinematic 720p videos — complete with expressive motion, dynamic lighting, and natural camera movement. Built for creators, storytellers, and developers, Vidu delivers smooth, detailed, and visually coherent motion sequences directly from text input.
prompt — describe your scene (e.g., “A cat walking through a neon-lit alley at night”).
movement_amplitude — control motion strength:
auto (default): model decides best motion level.
small: subtle, gentle movements (good for portraits or still scenes).
medium: balanced camera and subject motion.
large: cinematic, dramatic, or action-heavy movement.
seed — set for reproducible results (leave blank for random).
| Resolution | Duration | Cost per Clip |
|---|---|---|
| 720p | 4s | $0.20 |
Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/vidu/text-to-video with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Text To Video below.
set -euo pipefail
: "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"
REQUEST_BODY=$(cat <<'JSON'
{
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"movement_amplitude": "auto"
}
JSON
)
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/vidu/text-to-video" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $WAVESPEED_API_KEY" \
-d "$REQUEST_BODY")
TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL=$(printf '%s' "$TASK" | jq -r '.urls.get // empty')
if [ -z "$RESULT_URL" ]; then
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
fi
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
-H "Authorization: Bearer $WAVESPEED_API_KEY")
RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
case "$STATUS" in
completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
failed|cancelled|timeout) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
created|processing) sleep 2 ;;
*) printf 'Unexpected status: %s
' "$STATUS" >&2; exit 1 ;;
esac
doneconst submitUrl = "https://api.wavespeed.ai/api/v3/vidu/text-to-video";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');
async function requestJson(url, options = {}) {
const response = await fetch(url, options);
if (!response.ok) throw new Error(await response.text());
return response.json();
}
// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
method: "POST",
headers: {
"Authorization": `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"movement_amplitude": "auto"
}),
});
const task = body.data ?? body;
if (!task.id) throw new Error("Submission response did not contain a prediction id");
const resultUrl = task.urls?.get ||
`https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;
// 2. Poll until the prediction finishes.
while (true) {
const resultBody = await requestJson(resultUrl, {
headers: { "Authorization": `Bearer ${apiKey}` },
});
const result = resultBody.data ?? resultBody;
if (result.status === "completed") {
console.log(result.outputs);
break;
}
if (["failed", "cancelled", "timeout"].includes(result.status)) throw new Error(JSON.stringify(result));
if (!["created", "processing"].includes(result.status)) throw new Error("Unexpected status: " + result.status);
await new Promise(resolve => setTimeout(resolve, 2000));
}import json
import os
import time
from urllib.request import Request, urlopen
api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"movement_amplitude": "auto"
}
def request_json(url, data=None):
request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
with urlopen(request) as response:
return json.load(response)
# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/vidu/text-to-video", json.dumps(payload).encode())
task = body.get("data", body)
if not task.get("id"):
raise RuntimeError("Submission response did not contain a prediction id")
result_url = task.get("urls", {}).get("get") or f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"
# 2. Poll until the prediction finishes.
while True:
result_body = request_json(result_url)
result = result_body.get("data", result_body)
status = result.get("status")
if status == "completed":
print(result.get("outputs", []))
break
if status in {"failed", "cancelled", "timeout"}:
raise RuntimeError(result)
if status not in {"created", "processing"}:
raise RuntimeError(f"Unexpected status: {status}")
time.sleep(2)Text To Video is a Vidu model for video generation, exposed as a REST API on WaveSpeedAI. Vidu Text to Video converts text prompts into high-quality 720p videos with exceptional visual fidelity and diverse motion dynamics. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. You can call it programmatically or try it from the playground above.
POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates production-oriented Python, JavaScript, and cURL examples with timeouts, transient-error handling, and safe GET retries. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/vidu/vidu-text-to-video.
Text To Video starts at $0.20 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.
Key inputs: `prompt`, `seed`, `movement_amplitude`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/vidu/vidu-text-to-video.
Median end-to-end generation time on WaveSpeedAI is around 267 seconds per request, based on recent successful runs. Queue time varies with global demand; live status is visible in the prediction record.
Commercial usage rights depend on the model's license, set by its provider (Vidu). The license summary appears on the model card above; see WaveSpeedAI's Terms of Service for platform-level conditions.