Pixverse Pixverse V6 Reference To Video
Playground
Try it on WaveSpeedAI!PixVerse V6 Reference to Video creates high-quality videos from prompts, up to 10 reference images, and up to 2 reference videos. Use image references for subject, product, character, and scene consistency, and video references for motion, camera movement, action timing, and visual style. Supports optional audio and up to 1080P output. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Features
PixVerse V6 Reference-to-Video creates high-quality videos from a text prompt and visual references. Provide up to 10 reference images and up to 2 reference videos, then describe how the subjects, scene, action, camera movement, and visual style should be combined into the final clip.
This model is designed for Fusion-style generation: it can preserve the identity or look of reference subjects, borrow scene or motion cues from uploaded references, and use your prompt to compose them into a coherent video. It supports multiple resolution tiers up to 1080p and optional audio generation.
- Need to generate without references? Try PixVerse V6 Text-to-Video
- Need to animate a single image? Try PixVerse V6 Image-to-Video
Why Choose This?
-
Image and video reference control
Use reference images for subject, character, product, background, or style consistency, and use video references for motion, camera movement, scene rhythm, or visual style. -
V6 Fusion generation
The model can understand subjects, actions, scenes, camera moves, and style cues from reference media, then modify or recombine them according to the prompt. -
Up to 10 images and 2 videos
Add multiple image references for richer composition, plus up to 2 reference videos when you need stronger motion or scene guidance. -
Optional audio generation
Enablegenerate_audio_switchto create synchronized audio for the generated video. -
Flexible resolution choices
Chooseauto,360p,540p,720p, or1080pdepending on preview speed, quality target, and publishing needs.
Parameters
| Parameter | Required | Description |
|---|---|---|
| prompt | Yes | Describe the final video, including subject behavior, scene details, motion, camera movement, and style. When using named references in your prompt, make the relationship between each reference and the scene clear. |
| images | No | Reference image URLs. Supports up to 10 images. Use these for subject identity, character appearance, product details, background, costume, object, or style references. |
| videos | No | Reference video URLs. Supports up to 2 videos per request. The total duration of all reference videos must not exceed 15 seconds. Use these for action, motion, camera movement, scene rhythm, or video style references. |
| duration | No | Target clip length in seconds when no reference video is provided. Default: 5. In the current model schema, this parameter accepts 1–10. When videos are provided, set duration to 0; the output duration follows the reference video duration. |
| resolution | No | Output resolution: auto, 360p, 540p, 720p, or 1080p. Default: 720p. |
| generate_audio_switch | No | Whether to generate synchronized audio for the output video. Default: false. |
How to Use
- Write your prompt — Describe who appears, what happens, where it happens, and how the camera should move.
- Add reference images optional — Use
imageswhen visual identity matters, such as the same person, product, outfit, prop, background, or style. - Add reference videos optional — Use
videoswhen motion matters, such as dance movement, handheld camera feel, product rotation, action timing, or scene pacing. - Set duration — Choose the target duration when no reference video is provided. When using video references, keep the combined reference-video duration within 15 seconds and set
durationto0. - Choose resolution — Use
360por540pfor faster iteration, then switch to720por1080pfor higher-quality output. - Configure audio optional — Enable
generate_audio_switchwhen the scene benefits from ambience, impact sounds, crowd noise, weather, or other synchronized audio. - Submit — Generate the final reference-guided video.
Pricing
Pricing is calculated per generated second. Billed duration is rounded up to the next whole second, with a minimum of 5 seconds and a maximum of 15 seconds. When videos are provided, pricing uses the video-reference rate. auto is charged at the same rate as 360p.
Without Video References
| Resolution | No Audio | With Audio |
|---|---|---|
| auto / 360p | $0.025/s | $0.035/s |
| 540p | $0.035/s | $0.045/s |
| 720p | $0.045/s | $0.060/s |
| 1080p | $0.090/s | $0.115/s |
With Video References
| Resolution | No Audio | With Audio |
|---|---|---|
| auto / 360p | $0.050/s | $0.070/s |
| 540p | $0.070/s | $0.090/s |
| 720p | $0.090/s | $0.120/s |
| 1080p | $0.180/s | $0.230/s |
Example Costs
| Scenario | Cost |
|---|---|
| 10s, 720p, no audio, without video references | $0.45 |
| 10s, 720p, with audio, without video references | $0.60 |
| 6s reference video, 1080p, no audio | $1.08 |
| 6s reference video, 1080p, with audio | $1.38 |
Best Use Cases
- Reference-guided social videos — Combine character, product, or background references into short-form clips.
- Motion imitation — Use reference videos to guide dance, gesture, camera movement, or action pacing.
- Product and brand content — Keep products visually consistent while generating lifestyle or advertising scenes.
- Narrative clips — Build story moments with stable characters, props, backgrounds, and directed camera movement.
- Style transfer through references — Borrow lighting, composition, or motion mood from a reference video while changing the subject or scene.
Pro Tips
- Keep reference media clear, relevant, and visually uncluttered.
- Use fewer references when you need strict control over one subject.
- Use more references when you need a richer composed scene.
- Describe which reference should control which part of the result, such as subject, outfit, background, pose, action, or camera movement.
- For video references, mention whether you want to preserve motion, copy camera movement, recreate the scene structure, or replace the subject.
- Avoid conflicting references, such as two different backgrounds or incompatible character looks, unless the prompt explains how to combine them.
Notes
- At least one reference image or reference video is required. Provide
images,videos, or both. - When using video references, set
durationto0. The output duration will automatically match the longest reference video. - Reference names used in the prompt should not contain spaces. Use names like
character01,product_ref, ormotionRefinstead of names with spaces.
Related Models
- PixVerse V6 Text-to-Video — Generate video from a text prompt only.
- PixVerse V6 Image-to-Video — Animate a single reference image into video.
## Authentication
For authentication details, please refer to the [Authentication Guide](/api-authentication).
## API Endpoints
### Submit Task & Query Result
<ApiTabs submitUrl={model.submitUrl} resultUrl={model.resultUrl} payload={model.defaultValues} />
## Parameters
### Task Submission Parameters
#### Request Parameters
<RequestParams params={model.params} />
#### Response Parameters
<SubmitResponse />
#### Result Request Parameters
| Parameter | Type | Required | Default | Description |
|-----------|------|----------|---------|-------------|
| id | string | Yes | - | Task ID |
#### Result Response Parameters
| Parameter | Type | Description |
|-----------|------|-------------|
| code | integer | HTTP status code (e.g., 200 for success) |
| message | string | Status message (e.g., "success") |
| data | object | The prediction data object containing all details |
| data.id | string | Unique identifier for the prediction |
| data.model | string | Model ID used for the prediction |
| data.outputs | array<string \| object> | Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model. |
| data.urls | object | Object containing related API endpoints |
| data.urls.get | string | URL to poll for the prediction result |
| data.status | string | Status: `created`, `processing`, `completed`, or `failed` |
| data.created_at | string | ISO timestamp of when the request was created |
| data.error | string | Error message (empty if no error occurred) |
| data.timings | object | Object containing timing details |
| data.timings.inference | integer | Inference time in milliseconds |