Alibaba Wan 2.7 Reference To Video API Documentation

Alibaba Wan 2.7 Reference To Video API Documentation

Playground

Try it on WaveSpeedAI!

WAN 2.7 Reference-to-Video turns character, prop, or scene references from images or videos into new video shots with preserved identity, style, and layout plus smooth, coherent motion. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

Features

Wan 2.7 Reference-to-Video generates new video scenes guided by reference videos and an optional reference image, maintaining consistent characters, styles, and visual identity. Upload one or more reference videos, describe the scene you want, and the model produces a coherent, character-consistent video that brings your references into a new context.


Why Choose This?

  • Multi-video reference support Upload multiple reference videos to combine characters or visual elements from different sources into a single new scene.

  • Character-consistent generation The model preserves the identity, appearance, and style of characters from your reference videos throughout the generated clip.

  • Optional reference image Provide an additional still image to further guide the visual composition or introduce a new element.

  • Negative prompt support Specify what you don’t want in the output for more precise scene control.

  • Prompt expansion Enable enable_prompt_expansion to let the model automatically enrich and optimize your prompt before generation.

  • Resolution options Generate at 720p or 1080p to match your delivery requirements.


Parameters

ParameterRequiredDescription
reference_imagesNoReference image URLs (max 5). Combined with videos, total must be 1-5.
videosYesOne or more reference videos. Click Add Item to include additional videos.
promptYesText description of the desired scene and action. Reference characters as “Video 1”, “Video 2” etc.
imageNoOptional reference image to supplement the video references.
negative_promptNoElements to exclude from the generated video.
resolutionNoOutput resolution: 720p (default) or 1080p.
aspect_ratioNoOutput aspect ratio. Default: 16:9.
durationNoClip length in seconds. Default: 5.
enable_prompt_expansionNoEnable automatic prompt optimization before generation. Default: off.
seedNoRandom seed for reproducible results. Use -1 for a random seed.

How to Use

  1. Upload your reference videos — provide one or more source videos via URL or drag-and-drop. Click Add Item to add more.
  2. Write your prompt — describe the new scene, referencing characters by position (e.g., “The characters in Video 1 and Video 2 are sitting in front of the TV and playing video games together.”).
  3. Upload reference image (optional) — provide a still image to supplement the visual references.
  4. Add negative prompt (optional) — specify elements you want to exclude from the output.
  5. Select resolution — 720p for standard output, 1080p for higher-quality results.
  6. Select aspect ratio — choose the format that fits your target platform.
  7. Set duration — choose your desired clip length in seconds.
  8. Enable prompt expansion (optional) — let the model automatically enrich your prompt before generation.
  9. Set seed (optional) — fix the seed to reproduce a specific result in future runs.
  10. Submit — generate, preview, and download your video.

Prompt Indexing Rule

Refer to reference media as Image 1, Image 2, … and Video 1, Video 2, … (capitalized, with a space). Images and videos are counted separately, each in the order you upload them. For example, with 2 videos and 2 reference images:

  • Video 1 = first video
  • Video 2 = second video
  • Image 1 = first reference image
  • Image 2 = second reference image

Example prompt: "Video 2 holds Image 1 and plays a gentle folk song in a cafe, Video 1 smiles while watching Video 2"


Pricing

Duration720p1080p
5s$1.00$1.60
10s$1.50$2.40
15s$2.00$3.20

Billing Rules

  • 720p: base rate + fixed reference processing cost
  • 1080p: 1.6× the 720p cost
  • Pricing includes a fixed overhead for reference video processing in addition to the selected duration

Best Use Cases

  • Character-Driven Storytelling — Place characters from multiple reference videos into entirely new scenarios.
  • Fan Content & IP Crossovers — Combine characters from different sources into a single coherent scene.
  • Marketing & Brand Video — Generate new scenes featuring consistent brand characters or spokespeople from reference footage.
  • Creative Concepting — Rapidly prototype multi-character scenes for pitching and storyboarding.
  • Social Media Content — Create novel, character-consistent short-form video from existing footage.

Pro Tips

  • Use “Video 1”, “Video 2” etc. in your prompt to refer to specific reference videos in order.
  • The more distinct and clear each reference video is, the better the character consistency in the output.
  • Use negative_prompt to prevent unintended blending of visual styles between reference videos.
  • Enable prompt expansion for shorter or less detailed prompts to get richer output automatically.
  • Start with 720p to test your scene composition before committing to a 1080p final render.

Notes

  • Both videos and prompt are required fields; all other parameters are optional.
  • Ensure video and image URLs are publicly accessible if using links rather than direct uploads.
  • Please ensure your content complies with usage policies.
  • At least one reference_image or reference_video is required.
  • The total number of reference_image and reference_video items must not exceed 5.
  • Subject reference images or videos should contain only one main subject.
  • reference_video can provide subject and voice-timbre reference; empty-scene videos are not recommended.

Authentication

For authentication details, please refer to the Authentication Guide.

API Endpoints

Submit Task & Query Result

set -euo pipefail

export WAVESPEED_API_KEY="your-api-key"

REQUEST_BODY=$(cat <<'JSON'
{
  "prompt": "A cinematic ocean wave at sunrise, highly detailed",
  "resolution": "720p",
  "aspect_ratio": "16:9",
  "duration": 5,
  "enable_prompt_expansion": false
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/alibaba/wan-2.7/reference-to-video" \
  -H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
  -H "Content-Type: application/json" \
  -d "${REQUEST_BODY}")

TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body \
    "${RESULT_URL}" \
    -H "Authorization: Bearer ${WAVESPEED_API_KEY}")
  RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
  STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')

  case "${STATUS}" in
    completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
    failed|cancelled|timeout|deleted) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
    *) sleep 2 ;;
  esac
done

Parameters

Task Submission Parameters

Request Parameters

ParameterTypeRequiredDefaultRangeDescription
promptstringYes-The positive prompt for the generation. Refer to reference media as Image 1, Image 2, ... and Video 1, Video 2, ...; images and videos are counted separately, each in array order. E.g. with 2 videos and 1 image: 'Video 1 smiles at Video 2, Video 2 holds Image 1 in a cafe'.
imagestringNo-URL to a single reference image.
videosarray<string>No-0 ~ 5 itemsArray of reference video URLs (max 5). Combined count of reference_images and videos must be between 1 and 5. Refer to them in the prompt as Video 1, Video 2, ... in array order.
reference_imagesarray<string>No-0 ~ 5 itemsArray of reference image URLs (max 5). Combined count of reference_images and videos must be between 1 and 5. Refer to them in the prompt as Image 1, Image 2, ... in array order; images are counted separately from videos.
negative_promptstringNo-The negative prompt for the generation.
resolutionstringNo720p720p, 1080pThe resolution of the generated video.
aspect_ratiostringNo16:916:9, 9:16, 1:1, 4:3, 3:4The aspect ratio of the generated video.
durationintegerNo52 ~ 10The duration of the generated media in seconds (2-10s).
enable_prompt_expansionbooleanNofalse-If set to true, the prompt optimizer will be enabled.
seedintegerNo--The random seed to use for the generation. -1 means a random seed will be used.

Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
data.idstringUnique identifier for the prediction, Task Id
data.modelstringModel ID used for the prediction
data.outputsarrayOutput values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed)
data.urlsobjectObject containing related API endpoints
data.statusstringTask status. completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses.
data.created_atstringISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”)
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds

Result Request Parameters

ParameterTypeRequiredDefaultDescription
idstringYes-Task ID

Result Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
dataobjectThe prediction data object containing all details
data.idstringUnique identifier for the prediction
data.modelstringModel ID used for the prediction
data.outputsarray<string | object>Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model.
data.urlsobjectObject containing related API endpoints
data.statusstringStatus: completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses
data.created_atstringISO timestamp of when the request was created
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds
© 2026 WaveSpeedAI. All rights reserved.