Minimax H3 Reference To Video API Documentation

Minimax H3 Reference To Video API Documentation

Playground

Try it on WaveSpeedAI!

MiniMax H3 Reference to Video generates coherent 2K videos from natural-language prompts and multimodal references, including images, videos, and audio, guiding subject consistency, motion, timing, visual style, and scene continuity. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Features

MiniMax H3 Reference-to-Video generates high-resolution videos from a text prompt and reference media. Provide at least one reference image or video, optionally add reference audio, then describe the desired scene, interaction, camera movement, and visual style to create a 2k video output.


Why Choose This?

  • Reference-guided video generation
    Generate videos using reference images or videos to guide the final result.

  • Image and video reference support
    Use reference images for visual style, characters, objects, or scene composition, and reference videos for motion or interaction guidance.

  • Optional reference audio
    Add reference audio when audio guidance is needed, together with image or video references.

  • 2K video output
    Create high-resolution video outputs with the fixed 2k resolution tier.

  • Flexible aspect ratios
    Supports wide, landscape, square, portrait, and vertical formats including 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16.

  • Selectable duration
    Generate videos from 4 to 15 seconds.


Parameters

ParameterRequiredDescription
promptYesText description of the desired video scene and interactions. Minimum length: 1 character.
reference_imagesNoReference image URLs. Supports up to 9 images. At least one reference image or video is required.
reference_videosNoReference video URLs (up to 3). Each clip is normalized to 2-15 seconds and the combined reference-video duration is capped at 15 seconds.
reference_audiosNoOptional reference audio URLs. Supports up to 3 audio files. Audio cannot be provided alone.
aspect_ratioNoOutput aspect ratio: 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16.
resolutionNoOutput video resolution. Supported values: 768p or 2k.
durationNoOutput video duration in seconds. Supported values: 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15.

How to Use

  1. Write your prompt — Describe the scene, subject interaction, motion, camera movement, lighting, mood, and visual style.
  2. Add reference images or videos — Provide at least one reference image or video to guide the generated result.
  3. Add reference audio optional — Upload reference audio only when image or video references are also provided.
  4. Choose aspect ratio — Select the layout that matches your target format.
  5. Set duration — Choose a video length from 4 to 15 seconds.
  6. Submit — Generate the video and retrieve the output URL.

Pricing

Pricing includes generated output duration and normalized reference-video duration.

ResolutionPer billed second
768p$0.10
2K$0.14

Each reference video is normalized to at least 2 seconds and capped at 15 seconds. The combined input is capped at 15 seconds and rounded up to a whole second for billing. The first 5 reference images and supported audio inputs are free; each additional reference image costs $0.05.

Example: 5 seconds of normalized reference video plus 5 seconds of output costs $1.00 at 768p or $1.40 at 2K.


Best Use Cases

  • Reference-guided cinematic videos — Generate high-resolution videos from visual references and prompt instructions.
  • Character and subject consistency — Use reference images to guide characters, products, objects, or visual identity.
  • Motion and interaction guidance — Use reference videos to guide movement, pacing, or interaction style.
  • Commercial and campaign videos — Create polished video concepts from product, brand, or scene references.
  • Story and scene prototyping — Turn reference media and written ideas into short 2K motion clips.
  • Visual concept development — Explore camera movement, lighting, and composition with stronger reference control.

Pro Tips

  • Use clear reference images when subject identity, product details, or style consistency matters.
  • Use reference videos when motion, interaction, or camera movement is important.
  • Keep reference media visually relevant to the final video you want.
  • Make the relationship between the prompt and references clear.
  • Use 16:9 for widescreen video, 9:16 for vertical mobile content, and 1:1 for square layouts.
  • Use shorter durations for quick prompt testing and longer durations when the scene needs more time to develop.
  • Keep the prompt focused on visible action, scene progression, and interaction details.

Authentication

For authentication details, please refer to the Authentication Guide.

API Endpoints

Submit Task & Query Result

set -euo pipefail

export WAVESPEED_API_KEY="your-api-key"

REQUEST_BODY=$(cat <<'JSON'
{
  "prompt": "A cinematic ocean wave at sunrise, highly detailed",
  "aspect_ratio": "16:9",
  "resolution": "768p",
  "duration": 5
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/minimax/h3/reference-to-video" \
  -H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
  -H "Content-Type: application/json" \
  -d "${REQUEST_BODY}")

TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body \
    "${RESULT_URL}" \
    -H "Authorization: Bearer ${WAVESPEED_API_KEY}")
  RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
  STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')

  case "${STATUS}" in
    completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
    failed|cancelled|timeout|deleted) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
    *) sleep 2 ;;
  esac
done

Parameters

Task Submission Parameters

Request Parameters

ParameterTypeRequiredDefaultRangeDescription
promptstringYes-Text description of the desired video scene and interactions.
reference_imagesarray<string>No-0 ~ 9 itemsReference image URLs. At least one reference image or video is required.
reference_videosarray<string>No-0 ~ 3 itemsReference video URLs (up to 3). Each clip is normalized to 2-15 seconds and the combined duration is capped at 15 seconds.
reference_audiosarray<string>No-0 ~ 3 itemsOptional reference audio URLs. Audio cannot be provided alone.
aspect_ratiostringNo16:921:9, 16:9, 4:3, 1:1, 3:4, 9:16Output aspect ratio.
resolutionstringNo768p768p, 2kOutput video resolution.
durationintegerNo54, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15Output video duration in seconds.

Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
data.idstringUnique identifier for the prediction, Task Id
data.modelstringModel ID used for the prediction
data.outputsarrayOutput values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed)
data.urlsobjectObject containing related API endpoints
data.statusstringTask status. completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses.
data.created_atstringISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”)
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds

Result Request Parameters

ParameterTypeRequiredDefaultDescription
idstringYes-Task ID

Result Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
dataobjectThe prediction data object containing all details
data.idstringUnique identifier for the prediction
data.modelstringModel ID used for the prediction
data.outputsarray<string | object>Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model.
data.urlsobjectObject containing related API endpoints
data.statusstringStatus: completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses
data.created_atstringISO timestamp of when the request was created
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds
© 2026 WaveSpeedAI. All rights reserved.