Minimax H3 Reference To Video

Minimax H3 Reference To Video

Playground

Try it on WaveSpeedAI!

MiniMax H3 Reference to Video generates coherent 2K videos from natural-language prompts and multimodal references, including images, videos, and audio, guiding subject consistency, motion, timing, visual style, and scene continuity. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Features

MiniMax H3 Reference-to-Video generates high-resolution videos from a text prompt and reference media. Provide at least one reference image or video, optionally add reference audio, then describe the desired scene, interaction, camera movement, and visual style to create a 2k video output.


Why Choose This?

  • Reference-guided video generation
    Generate videos using reference images or videos to guide the final result.

  • Image and video reference support
    Use reference images for visual style, characters, objects, or scene composition, and reference videos for motion or interaction guidance.

  • Optional reference audio
    Add reference audio when audio guidance is needed, together with image or video references.

  • 2K video output
    Create high-resolution video outputs with the fixed 2k resolution tier.

  • Flexible aspect ratios
    Supports wide, landscape, square, portrait, and vertical formats including 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16.

  • Selectable duration
    Generate videos from 5 to 15 seconds.


Parameters

ParameterRequiredDescription
promptYesText description of the desired video scene and interactions. Minimum length: 1 character. Maximum length: 4000 characters.
reference_imagesNoReference image URLs. Supports up to 9 images. At least one reference image or video is required.
reference_videosNoReference video URLs. Supports up to 3 videos. At least one reference image or video is required.
reference_audiosNoOptional reference audio URLs. Supports up to 3 audio files. Audio cannot be provided alone.
aspect_ratioNoOutput aspect ratio: 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16.
resolutionNoOutput video resolution. Supported value: 2k.
durationNoOutput video duration in seconds. Supported values: 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15.

How to Use

  1. Write your prompt — Describe the scene, subject interaction, motion, camera movement, lighting, mood, and visual style.
  2. Add reference images or videos — Provide at least one reference image or video to guide the generated result.
  3. Add reference audio optional — Upload reference audio only when image or video references are also provided.
  4. Choose aspect ratio — Select the layout that matches your target format.
  5. Set duration — Choose a video length from 5 to 15 seconds.
  6. Submit — Generate the video and retrieve the output URL.

Pricing

Generated 2K video costs $0.13 per output second.

Output DurationOutput Price
5s$0.65
10s$1.30
15s$1.95

Billing Rules

  • Reference inputs: the first 5 images and all supported audio inputs are free. Each image after the first 5 costs $0.04. Reference videos cost $0.13 per second based on their combined duration, with up to 3 videos and 15 seconds total per request.

Best Use Cases

  • Reference-guided cinematic videos — Generate high-resolution videos from visual references and prompt instructions.
  • Character and subject consistency — Use reference images to guide characters, products, objects, or visual identity.
  • Motion and interaction guidance — Use reference videos to guide movement, pacing, or interaction style.
  • Commercial and campaign videos — Create polished video concepts from product, brand, or scene references.
  • Story and scene prototyping — Turn reference media and written ideas into short 2K motion clips.
  • Visual concept development — Explore camera movement, lighting, and composition with stronger reference control.

Pro Tips

  • Use clear reference images when subject identity, product details, or style consistency matters.
  • Use reference videos when motion, interaction, or camera movement is important.
  • Keep reference media visually relevant to the final video you want.
  • Make the relationship between the prompt and references clear.
  • Use 16:9 for widescreen video, 9:16 for vertical mobile content, and 1:1 for square layouts.
  • Use shorter durations for quick prompt testing and longer durations when the scene needs more time to develop.
  • Keep the prompt focused on visible action, scene progression, and interaction details.

Authentication

For authentication details, please refer to the Authentication Guide.

API Endpoints

Submit Task & Query Result

set -euo pipefail

export WAVESPEED_API_KEY="your-api-key"

REQUEST_BODY=$(cat <<'JSON'
{
  "prompt": "A cinematic ocean wave at sunrise, highly detailed",
  "aspect_ratio": "21:9",
  "resolution": "2k",
  "duration": 5
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/minimax/h3/reference-to-video" \
  -H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
  -H "Content-Type: application/json" \
  -d "${REQUEST_BODY}")

TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL=$(printf '%s' "${TASK}" | jq -r '.urls.get // empty')
if [ -z "${RESULT_URL}" ]; then RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"; fi

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body \
    "${RESULT_URL}" \
    -H "Authorization: Bearer ${WAVESPEED_API_KEY}")
  RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
  STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')

  case "${STATUS}" in
    completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
    failed|cancelled|timeout) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
    created|processing) sleep 2 ;;
    *) printf 'Unexpected status: %s
' "${STATUS}" >&2; exit 1 ;;
  esac
done

Parameters

Task Submission Parameters

Request Parameters

ParameterTypeRequiredDefaultRangeDescription
promptstringYes-Text description of the desired video scene and interactions.
reference_imagesarray<string>No-0 ~ 9 itemsReference image URLs. At least one reference image or video is required.
reference_videosarray<string>No-0 ~ 3 itemsReference video URLs. At least one reference image or video is required.
reference_audiosarray<string>No-0 ~ 3 itemsOptional reference audio URLs. Audio cannot be provided alone.
aspect_ratiostringNo-21:9, 16:9, 4:3, 1:1, 3:4, 9:16Output aspect ratio.
resolutionstringNo2k2kOutput video resolution.
durationintegerNo55, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15Output video duration in seconds.

Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
data.idstringUnique identifier for the prediction, Task Id
data.modelstringModel ID used for the prediction
data.outputsarrayOutput values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed)
data.urlsobjectObject containing related API endpoints
data.urls.getstringURL to retrieve the prediction result
data.statusstringStatus of the task: created, processing, completed, or failed
data.created_atstringISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”)
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds

Result Request Parameters

ParameterTypeRequiredDefaultDescription
idstringYes-Task ID

Result Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
dataobjectThe prediction data object containing all details
data.idstringUnique identifier for the prediction
data.modelstringModel ID used for the prediction
data.outputsarray<string | object>Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model.
data.urlsobjectObject containing related API endpoints
data.urls.getstringURL to poll for the prediction result
data.statusstringStatus: created, processing, completed, or failed
data.created_atstringISO timestamp of when the request was created
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds
© 2026 WaveSpeedAI. All rights reserved.