Seedream 5.0 Pro 정식 출시 | 이미지 생성기에서 사용해보기 →
/탐색/PixVerse/Pixverse V6/Reference To Video

PixVerse V6 Reference to Video

pixverse /

PixVerse V6 Reference to Video creates high-quality videos from prompts, up to 10 reference images, and up to 2 reference videos. Use image references for subject, product, character, and scene consistency, and video references for motion, camera movement, action timing, and visual style. Supports optional audio and up to 1080P output. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

image-to-video
입력

대기 중

$1실행당

다음:

관련 모델

README

PixVerse V6 Reference-to-Video

PixVerse V6 Reference-to-Video creates high-quality videos from a text prompt and visual references. Provide up to 10 reference images and up to 2 reference videos, then describe how the subjects, scene, action, camera movement, and visual style should be combined into the final clip.

This model is designed for Fusion-style generation: it can preserve the identity or look of reference subjects, borrow scene or motion cues from uploaded references, and use your prompt to compose them into a coherent video. It supports multiple resolution tiers up to 1080p and optional audio generation.

Why Choose This?

  • Image and video reference control
    Use reference images for subject, character, product, background, or style consistency, and use video references for motion, camera movement, scene rhythm, or visual style.

  • V6 Fusion generation
    The model can understand subjects, actions, scenes, camera moves, and style cues from reference media, then modify or recombine them according to the prompt.

  • Up to 10 images and 2 videos
    Add multiple image references for richer composition, plus up to 2 reference videos when you need stronger motion or scene guidance.

  • Optional audio generation
    Enable generate_audio_switch to create synchronized audio for the generated video.

  • Flexible resolution choices
    Choose auto, 360p, 540p, 720p, or 1080p depending on preview speed, quality target, and publishing needs.

Parameters

ParameterRequiredDescription
promptYesDescribe the final video, including subject behavior, scene details, motion, camera movement, and style. When using named references in your prompt, make the relationship between each reference and the scene clear.
imagesNoReference image URLs. Supports up to 10 images. Use these for subject identity, character appearance, product details, background, costume, object, or style references.
videosNoReference video URLs. Supports up to 2 videos per request. The total duration of all reference videos must not exceed 15 seconds. Use these for action, motion, camera movement, scene rhythm, or video style references.
durationNoTarget clip length in seconds when no reference video is provided. Default: 5. In the current model schema, this parameter accepts 1–10. When videos are provided, set duration to 0; the output duration follows the reference video duration.
resolutionNoOutput resolution: auto, 360p, 540p, 720p, or 1080p. Default: 720p.
generate_audio_switchNoWhether to generate synchronized audio for the output video. Default: false.

How to Use

  1. Write your prompt — Describe who appears, what happens, where it happens, and how the camera should move.
  2. Add reference images optional — Use images when visual identity matters, such as the same person, product, outfit, prop, background, or style.
  3. Add reference videos optional — Use videos when motion matters, such as dance movement, handheld camera feel, product rotation, action timing, or scene pacing.
  4. Set duration — Choose the target duration when no reference video is provided. When using video references, keep the combined reference-video duration within 15 seconds and set duration to 0.
  5. Choose resolution — Use 360p or 540p for faster iteration, then switch to 720p or 1080p for higher-quality output.
  6. Configure audio optional — Enable generate_audio_switch when the scene benefits from ambience, impact sounds, crowd noise, weather, or other synchronized audio.
  7. Submit — Generate the final reference-guided video.

Pricing

Pricing is calculated per generated second. Billed duration is rounded up to the next whole second, with a minimum of 5 seconds and a maximum of 15 seconds. When videos are provided, pricing uses the video-reference rate. auto is charged at the same rate as 360p.

Without Video References

ResolutionNo AudioWith Audio
auto / 360p$0.025/s$0.035/s
540p$0.035/s$0.045/s
720p$0.045/s$0.060/s
1080p$0.090/s$0.115/s

With Video References

ResolutionNo AudioWith Audio
auto / 360p$0.050/s$0.070/s
540p$0.070/s$0.090/s
720p$0.090/s$0.120/s
1080p$0.180/s$0.230/s

Example Costs

ScenarioCost
10s, 720p, no audio, without video references$0.45
10s, 720p, with audio, without video references$0.60
6s reference video, 1080p, no audio$1.08
6s reference video, 1080p, with audio$1.38

Best Use Cases

  • Reference-guided social videos — Combine character, product, or background references into short-form clips.
  • Motion imitation — Use reference videos to guide dance, gesture, camera movement, or action pacing.
  • Product and brand content — Keep products visually consistent while generating lifestyle or advertising scenes.
  • Narrative clips — Build story moments with stable characters, props, backgrounds, and directed camera movement.
  • Style transfer through references — Borrow lighting, composition, or motion mood from a reference video while changing the subject or scene.

Pro Tips

  • Keep reference media clear, relevant, and visually uncluttered.
  • Use fewer references when you need strict control over one subject.
  • Use more references when you need a richer composed scene.
  • Describe which reference should control which part of the result, such as subject, outfit, background, pose, action, or camera movement.
  • For video references, mention whether you want to preserve motion, copy camera movement, recreate the scene structure, or replace the subject.
  • Avoid conflicting references, such as two different backgrounds or incompatible character looks, unless the prompt explains how to combine them.

Notes

  • At least one reference image or reference video is required. Provide images, videos, or both.
  • When using video references, set duration to 0. The output duration will automatically match the longest reference video.
  • Reference names used in the prompt should not contain spaces. Use names like character01, product_ref, or motionRef instead of names with spaces.

Related Models

참고:이 웹사이트는 제3자가 제공하는 AI 모델을 사용합니다.

Pixverse v6 Reference To Video API — Quick start

Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/pixverse/pixverse-v6/reference-to-video with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Pixverse v6 Reference To Video below.

HTTP example
set -euo pipefail

: "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"

REQUEST_BODY=$(cat <<'JSON'
{
    "prompt": "A cinematic shot of a city at sunset, soft golden light",
    "resolution": "720p",
    "duration": 5,
    "aspect_ratio": "auto",
    "generate_audio_switch": false
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/pixverse/pixverse-v6/reference-to-video" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $WAVESPEED_API_KEY" \
  -d "$REQUEST_BODY")

TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL=$(printf '%s' "$TASK" | jq -r '.urls.get // empty')
if [ -z "$RESULT_URL" ]; then
  RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
fi

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
    -H "Authorization: Bearer $WAVESPEED_API_KEY")
  RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
  STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
  case "$STATUS" in
    completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
    failed|cancelled|timeout) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
    created|processing) sleep 2 ;;
    *) printf 'Unexpected status: %s
' "$STATUS" >&2; exit 1 ;;
  esac
done
Node.js example
const submitUrl = "https://api.wavespeed.ai/api/v3/pixverse/pixverse-v6/reference-to-video";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');

async function requestJson(url, options = {}) {
  const response = await fetch(url, options);
  if (!response.ok) throw new Error(await response.text());
  return response.json();
}

// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${apiKey}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
        "prompt": "A cinematic shot of a city at sunset, soft golden light",
        "resolution": "720p",
        "duration": 5,
        "aspect_ratio": "auto",
        "generate_audio_switch": false
}),
});
const task = body.data ?? body;
if (!task.id) throw new Error("Submission response did not contain a prediction id");
const resultUrl = task.urls?.get ||
  `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;

// 2. Poll until the prediction finishes.
while (true) {
  const resultBody = await requestJson(resultUrl, {
    headers: { "Authorization": `Bearer ${apiKey}` },
  });
  const result = resultBody.data ?? resultBody;
  if (result.status === "completed") {
    console.log(result.outputs);
    break;
  }
  if (["failed", "cancelled", "timeout"].includes(result.status)) throw new Error(JSON.stringify(result));
  if (!["created", "processing"].includes(result.status)) throw new Error("Unexpected status: " + result.status);
  await new Promise(resolve => setTimeout(resolve, 2000));
}
Python example
import json
import os
import time
from urllib.request import Request, urlopen

api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
    "prompt": "A cinematic shot of a city at sunset, soft golden light",
    "resolution": "720p",
    "duration": 5,
    "aspect_ratio": "auto",
    "generate_audio_switch": False
}

def request_json(url, data=None):
    request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
    with urlopen(request) as response:
        return json.load(response)

# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/pixverse/pixverse-v6/reference-to-video", json.dumps(payload).encode())
task = body.get("data", body)
if not task.get("id"):
    raise RuntimeError("Submission response did not contain a prediction id")
result_url = task.get("urls", {}).get("get") or f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"

# 2. Poll until the prediction finishes.
while True:
    result_body = request_json(result_url)
    result = result_body.get("data", result_body)
    status = result.get("status")
    if status == "completed":
        print(result.get("outputs", []))
        break
    if status in {"failed", "cancelled", "timeout"}:
        raise RuntimeError(result)
    if status not in {"created", "processing"}:
        raise RuntimeError(f"Unexpected status: {status}")
    time.sleep(2)

Pixverse v6 Reference To Video API — Frequently asked questions

What is the Pixverse v6 Reference To Video API?

Pixverse v6 Reference To Video is a Pixverse model for video generation from images, exposed as a REST API on WaveSpeedAI. PixVerse V6 Reference to Video creates high-quality videos from prompts, up to 10 reference images, and up to 2 reference videos. Use image references for subject, product, character, and scene consistency, and video references for motion, camera movement, action timing, and visual style. Supports optional audio and up to 1080P output. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. You can call it programmatically or try it from the playground above.

How do I call the Pixverse v6 Reference To Video API?

POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates production-oriented Python, JavaScript, and cURL examples with timeouts, transient-error handling, and safe GET retries. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/pixverse/pixverse-pixverse-v6-reference-to-video.

How much does Pixverse v6 Reference To Video cost per run?

Pixverse v6 Reference To Video starts at $1.00 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.

What inputs does Pixverse v6 Reference To Video accept?

Key inputs: `prompt`, `images`, `aspect_ratio`, `resolution`, `duration`, `generate_audio_switch`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/pixverse/pixverse-pixverse-v6-reference-to-video.

How do I get started with the Pixverse v6 Reference To Video API?

Sign up for a free WaveSpeedAI account to claim starter credits, copy your API key from /accesskey, then call the endpoint shown in the API tab of the playground. The playground also auto-generates a code sample in Python, JavaScript, or cURL for the parameters you've set.

Can I use Pixverse v6 Reference To Video outputs commercially?

Commercial usage rights depend on the model's license, set by its provider (Pixverse). The license summary appears on the model card above; see WaveSpeedAI's Terms of Service for platform-level conditions.

PixVerse V6 Reference to Video | WaveSpeedAI