Pixverse Pixverse V6 Reference To Video

Pixverse Pixverse V6 Reference To Video

Playground

Try it on WaveSpeedAI!

PixVerse V6 Reference to Video creates high-quality videos from prompts, up to 10 reference images, and up to 2 reference videos. Use image references for subject, product, character, and scene consistency, and video references for motion, camera movement, action timing, and visual style. Supports optional audio and up to 1080P output. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Features

PixVerse V6 Reference-to-Video creates high-quality videos from a text prompt and visual references. Provide up to 10 reference images and up to 2 reference videos, then describe how the subjects, scene, action, camera movement, and visual style should be combined into the final clip.

This model is designed for Fusion-style generation: it can preserve the identity or look of reference subjects, borrow scene or motion cues from uploaded references, and use your prompt to compose them into a coherent video. It supports multiple resolution tiers up to 1080p and optional audio generation.


Why Choose This?

  • Image and video reference control
    Use reference images for subject, character, product, background, or style consistency, and use video references for motion, camera movement, scene rhythm, or visual style.

  • V6 Fusion generation
    The model can understand subjects, actions, scenes, camera moves, and style cues from reference media, then modify or recombine them according to the prompt.

  • Up to 10 images and 2 videos
    Add multiple image references for richer composition, plus up to 2 reference videos when you need stronger motion or scene guidance.

  • Optional audio generation
    Enable generate_audio_switch to create synchronized audio for the generated video.

  • Flexible resolution choices
    Choose auto, 360p, 540p, 720p, or 1080p depending on preview speed, quality target, and publishing needs.


Parameters

ParameterRequiredDescription
promptYesDescribe the final video, including subject behavior, scene details, motion, camera movement, and style. When using named references in your prompt, make the relationship between each reference and the scene clear.
imagesNoReference image URLs. Supports up to 10 images. Use these for subject identity, character appearance, product details, background, costume, object, or style references.
videosNoReference video URLs. Supports up to 2 videos per request. The total duration of all reference videos must not exceed 15 seconds. Use these for action, motion, camera movement, scene rhythm, or video style references.
durationNoTarget clip length in seconds when no reference video is provided. Default: 5. In the current model schema, this parameter accepts 1–10. When videos are provided, set duration to 0; the output duration follows the reference video duration.
resolutionNoOutput resolution: auto, 360p, 540p, 720p, or 1080p. Default: 720p.
generate_audio_switchNoWhether to generate synchronized audio for the output video. Default: false.

How to Use

  1. Write your prompt — Describe who appears, what happens, where it happens, and how the camera should move.
  2. Add reference images optional — Use images when visual identity matters, such as the same person, product, outfit, prop, background, or style.
  3. Add reference videos optional — Use videos when motion matters, such as dance movement, handheld camera feel, product rotation, action timing, or scene pacing.
  4. Set duration — Choose the target duration when no reference video is provided. When using video references, keep the combined reference-video duration within 15 seconds and set duration to 0.
  5. Choose resolution — Use 360p or 540p for faster iteration, then switch to 720p or 1080p for higher-quality output.
  6. Configure audio optional — Enable generate_audio_switch when the scene benefits from ambience, impact sounds, crowd noise, weather, or other synchronized audio.
  7. Submit — Generate the final reference-guided video.

Pricing

Pricing is calculated per generated second. Billed duration is rounded up to the next whole second, with a minimum of 5 seconds and a maximum of 15 seconds. When videos are provided, pricing uses the video-reference rate. auto is charged at the same rate as 360p.

Without Video References

ResolutionNo AudioWith Audio
auto / 360p$0.025/s$0.035/s
540p$0.035/s$0.045/s
720p$0.045/s$0.060/s
1080p$0.090/s$0.115/s

With Video References

ResolutionNo AudioWith Audio
auto / 360p$0.050/s$0.070/s
540p$0.070/s$0.090/s
720p$0.090/s$0.120/s
1080p$0.180/s$0.230/s

Example Costs

ScenarioCost
10s, 720p, no audio, without video references$0.45
10s, 720p, with audio, without video references$0.60
6s reference video, 1080p, no audio$1.08
6s reference video, 1080p, with audio$1.38

Best Use Cases

  • Reference-guided social videos — Combine character, product, or background references into short-form clips.
  • Motion imitation — Use reference videos to guide dance, gesture, camera movement, or action pacing.
  • Product and brand content — Keep products visually consistent while generating lifestyle or advertising scenes.
  • Narrative clips — Build story moments with stable characters, props, backgrounds, and directed camera movement.
  • Style transfer through references — Borrow lighting, composition, or motion mood from a reference video while changing the subject or scene.

Pro Tips

  • Keep reference media clear, relevant, and visually uncluttered.
  • Use fewer references when you need strict control over one subject.
  • Use more references when you need a richer composed scene.
  • Describe which reference should control which part of the result, such as subject, outfit, background, pose, action, or camera movement.
  • For video references, mention whether you want to preserve motion, copy camera movement, recreate the scene structure, or replace the subject.
  • Avoid conflicting references, such as two different backgrounds or incompatible character looks, unless the prompt explains how to combine them.

Notes

  • At least one reference image or reference video is required. Provide images, videos, or both.
  • When using video references, set duration to 0. The output duration will automatically match the longest reference video.
  • Reference names used in the prompt should not contain spaces. Use names like character01, product_ref, or motionRef instead of names with spaces.


## Authentication

For authentication details, please refer to the [Authentication Guide](/api-authentication).

## API Endpoints

### Submit Task & Query Result

<ApiTabs submitUrl={model.submitUrl} resultUrl={model.resultUrl} payload={model.defaultValues} />

## Parameters

### Task Submission Parameters

#### Request Parameters

<RequestParams params={model.params} />

#### Response Parameters

<SubmitResponse />

#### Result Request Parameters

| Parameter | Type | Required | Default | Description |
|-----------|------|----------|---------|-------------|
| id | string | Yes | - | Task ID |

#### Result Response Parameters

| Parameter | Type | Description |
|-----------|------|-------------|
| code | integer | HTTP status code (e.g., 200 for success) |
| message | string | Status message (e.g., "success") |
| data | object | The prediction data object containing all details |
| data.id | string | Unique identifier for the prediction |
| data.model | string | Model ID used for the prediction |
| data.outputs | array&lt;string \| object&gt; | Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model. |
| data.urls | object | Object containing related API endpoints |
| data.urls.get | string | URL to poll for the prediction result |
| data.status | string | Status: `created`, `processing`, `completed`, or `failed` |
| data.created_at | string | ISO timestamp of when the request was created |
| data.error | string | Error message (empty if no error occurred) |
| data.timings | object | Object containing timing details |
| data.timings.inference | integer | Inference time in milliseconds |
© 2026 WaveSpeedAI. All rights reserved.