Kling O3 Omni 4K Video Reference generates 4K AI videos guided by an input video, text prompt, and optional reference images, supporting video-to-video creation with visual reference guidance, motion consistency, camera control, and high-quality cinematic output. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Idle
$2.31per run
Use the reference video as the motion and camera guide. Preserve the woman’s running rhythm, body movement, turning action, jump timing, movement direction, shot composition, and camera trajectory. Transform the runner into a young elven royal messenger with long silver-blonde hair, pointed ears, a fitted dark green leather tunic, layered forest-green cloak, brown riding boots, silver leaf-shaped accessories, and a small weathered messenger satchel. Replace the modern city street with an ancient Western fantasy mountain city at twilight, featuring weathered stone stairways, tall medieval towers, hanging banners, glowing lanterns, narrow arched passages, and distant mountains covered in mist. As she runs, her cloak and hair flow naturally behind her. During the jump, several loose parchment letters briefly move inside the open satchel. Add subtle warm lantern light contrasting against the cold blue twilight. Maintain the exact pacing and continuous camera movement of the reference video. Keep the character visually consistent throughout the sequence. Cinematic Western fantasy adventure, realistic materials, atmospheric mist, elegant production design, high-end fantasy film aesthetic.
Kling O3 Omni 4K Video Reference generates a new 4K video using an input video as a motion, appearance, or scene reference. Provide a reference video and a text prompt, then optionally add image references or element IDs for stronger visual guidance.
Video-reference generation
Use an input video as a reference for motion, appearance, subject behavior, or scene guidance.
4K video output
Generate high-resolution 4K video for higher-quality delivery workflows.
Prompt-guided control
Describe the new scene, subject behavior, camera movement, and visual treatment with text instructions.
Image reference support
Add optional image references when extra subject, object, or style guidance is needed.
Element reference support
Use element_list to reference reusable Kling elements by element_id.
Original sound preservation
Keep the original audio from the reference video with keep_original_sound.
| Parameter | Required | Description |
|---|---|---|
| prompt | Yes | Describe the new scene, subject behavior, camera movement, and visual treatment. |
| video | Yes | Reference video used for motion, appearance, or scene guidance. |
| images | No | Optional reference images. Supports up to 4 image URLs. Combined official reference limits still apply. |
| keep_original_sound | No | Keep the original audio from the reference video. Default: true. |
| aspect_ratio | No | Output video aspect ratio: 16:9, 9:16, or 1:1. Default: 16:9. |
| duration | No | Generated video duration in seconds. Range: 3–15. Default: 5. |
| element_list | No | Element reference list. Supports up to 3 items. Each item uses an element_id returned by kwaivgi/kling-elements or kwaivgi/kling-elements-advanced. |
images when extra subject, object, or style reference is needed.element_list when you want to guide generation with existing Kling element IDs.16:9, 9:16, or 1:1.3 to 15 seconds.keep_original_sound enabled when the reference video's audio should be preserved.Pricing is based on generated video duration.
| Billing Unit | Price |
|---|---|
| Per second | $0.462 |
| Per 5 seconds | $2.31 |
| Duration | Price |
|---|---|
| 3 seconds | $1.386 |
| 5 seconds | $2.31 |
| 10 seconds | $4.62 |
| 15 seconds | $6.93 |
element_list when you already have reusable Kling element IDs.element_list focused; too many unrelated elements can reduce control.keep_original_sound only when the input video's audio should be carried into the result.9:16 for vertical content, 16:9 for widescreen output, and 1:1 for square formats.prompt and video are required.images supports up to 4 reference images.element_list supports up to 3 element references.element_list item must include an element_id.element_id should come from kwaivgi/kling-elements or kwaivgi/kling-elements-advanced.duration must be between 3 and 15 seconds.keep_original_sound defaults to true.element_id references.Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/kwaivgi/kling-video-o3-4k/video-reference with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Kling Video O3 4k Video Reference below.
set -euo pipefail
: "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"
REQUEST_BODY=$(cat <<'JSON'
{
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
"keep_original_sound": true,
"aspect_ratio": "16:9",
"duration": 5
}
JSON
)
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/kwaivgi/kling-video-o3-4k/video-reference" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $WAVESPEED_API_KEY" \
-d "$REQUEST_BODY")
TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
-H "Authorization: Bearer $WAVESPEED_API_KEY")
RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
case "$STATUS" in
completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
failed|cancelled|timeout|deleted) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
*) sleep 2 ;;
esac
doneconst submitUrl = "https://api.wavespeed.ai/api/v3/kwaivgi/kling-video-o3-4k/video-reference";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');
async function requestJson(url, options = {}) {
const response = await fetch(url, options);
if (!response.ok) throw new Error(await response.text());
return response.json();
}
// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
method: "POST",
headers: {
"Authorization": `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
"keep_original_sound": true,
"aspect_ratio": "16:9",
"duration": 5
}),
});
const task = body.data ?? body;
if (!task.id) throw new Error("Submission response did not contain a prediction id");
const resultUrl = `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;
// 2. Poll until the prediction finishes.
while (true) {
const resultBody = await requestJson(resultUrl, {
headers: { "Authorization": `Bearer ${apiKey}` },
});
const result = resultBody.data ?? resultBody;
if (result.status === "completed") {
console.log(result.outputs);
break;
}
if (["failed", "cancelled", "timeout", "deleted"].includes(result.status)) throw new Error(JSON.stringify(result));
await new Promise(resolve => setTimeout(resolve, 2000));
}import json
import os
import time
from urllib.request import Request, urlopen
api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
"keep_original_sound": True,
"aspect_ratio": "16:9",
"duration": 5
}
def request_json(url, data=None):
request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
with urlopen(request) as response:
return json.load(response)
# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/kwaivgi/kling-video-o3-4k/video-reference", json.dumps(payload).encode())
task = body.get("data", body)
if not task.get("id"):
raise RuntimeError("Submission response did not contain a prediction id")
result_url = f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"
# 2. Poll until the prediction finishes.
while True:
result_body = request_json(result_url)
result = result_body.get("data", result_body)
status = result.get("status")
if status == "completed":
print(result.get("outputs", []))
break
if status in {"failed", "cancelled", "timeout", "deleted"}:
raise RuntimeError(result)
time.sleep(2)Kling Video O3 4k Video Reference is a Kuaishou model for video editing, exposed as a REST API on WaveSpeedAI. Kling O3 Omni 4K Video Reference generates 4K AI videos guided by an input video, text prompt, and optional reference images, supporting video-to-video creation with visual reference guidance, motion consistency, camera control, and high-quality cinematic output. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. You can call it programmatically or try it from the playground above.
POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates production-oriented Python, JavaScript, and cURL examples with timeouts, transient-error handling, and safe GET retries. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/kwaivgi/kwaivgi-kling-video-o3-4k-video-reference.
Kling Video O3 4k Video Reference starts at $2.31 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.
Key inputs: `prompt`, `images`, `video`, `aspect_ratio`, `duration`, `element_list`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/kwaivgi/kwaivgi-kling-video-o3-4k-video-reference.
Sign up for a free WaveSpeedAI account to claim starter credits, copy your API key from /accesskey, then call the endpoint shown in the API tab of the playground. The playground also auto-generates a code sample in Python, JavaScript, or cURL for the parameters you've set.
Commercial usage rights depend on the model's license, set by its provider (Kuaishou). The license summary appears on the model card above; see WaveSpeedAI's Terms of Service for platform-level conditions.