Minimax H3 Reference To Video LoRA API Documentation

Minimax H3 Reference To Video LoRA API Documentation

Playground

Try it on WaveSpeedAI!

MiniMax H3 Open Weights Reference to Video with custom LoRA support generates coherent 480P / 768P videos from prompts and multimodal references, guided by up to 9 reference images, 3 reference videos, and 3 reference audios, with native stereo audio and flexible reference-based video generation on WaveSpeedAI infrastructure. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Features

Run the open-weights edition of MiniMax H3 on WaveSpeedAI’s own GPU infrastructure. This endpoint is separate from the official minimax/h3 API: same model family, independently hosted, with its own 480p/768p resolutions and per-second pricing.

MiniMax H3 is an omni-modal video model that generates picture and native stereo audio in a single pass. In reference-to-video mode you guide the generation with up to 9 reference images, 3 reference videos, and 3 reference audio tracks, then tell the model — in the prompt — what each reference is for.


How to Use References

Refer to every input in the prompt with an angle-bracket tag: <Picture 1>, <Video 1>, <Audio 1>, and so on. The tags must be written exactly like that, with the brackets — plain text such as Picture 1 is treated as ordinary words, not a reference.

Assign a job to each reference. Explicit assignments work far better than leaving the model to guess:

> “Use the character from <Picture 1>, place them in the setting from <Picture 2>, and match the camera motion of <Video 1>.”

Numbering

  • Tags are numbered per type, in the order the inputs are provided: images are <Picture 1><Picture 9>, standalone audios are <Audio 1>…, videos are <Video 1><Video 3>.
  • A reference video’s own soundtrack is used automatically and occupies the earliest <Audio …> slots; any standalone reference_audios you provide are numbered after the video soundtracks. If you only need the video’s audio, refer to it via its <Video …> tag.

Rules

  • At least one reference image or video is required. Audio cannot be provided alone.
  • reference_videos is available only at 480p output resolution.
  • Reference video soundtracks are picked up automatically; total reference-video duration shares a 15-second budget and longer inputs are trimmed fairly.

How to Write a Great Prompt

Write the prompt as a timeline with a schedule, and weave the reference tags into it:

  1. Style & references — establish the look and say which reference drives identity, style, motion, or voice.
  2. Timeline — timestamped beats: “[0s-3s] <Picture 1> stands on the balcony from <Picture 2> … [3s-6s] turns toward camera and smiles …”.
  3. Camera — explicit movement, including when to hold: “slow orbit, no cuts”.
  4. Audio — an Audio: line with entrance cues; reference an input voice or track by tag if you want it matched.
  5. On-screen text — spell out readable words in quotes and add “no subtitles, do not add other text.”

Parameters

ParameterRequiredDescription
promptYesDesired video with references addressed as <Picture 1>, <Video 1>, <Audio 1>, plus an Audio: line for the soundtrack.
reference_imagesNoUp to 9 reference image URLs.
reference_videosNoUp to 3 reference video URLs. 480p output only. Their soundtracks are used automatically.
reference_audiosNoUp to 3 standalone reference audio URLs (each trimmed to 15s).
aspect_ratioNo16:9, 9:16, 1:1, 4:3, 3:4, 21:9, or 9:21. Default: 16:9.
resolutionNo480p (faster, lower cost) or 768p (native canvas). Default: 480p.
durationNoOutput length in seconds, 315. Default: 5.
seedNoFixed seed for reproducible output.
lorasNoUp to 3 LoRA weights, each {path, scale}; path is a LoRA file URL.

At least one of reference_images, reference_videos, or reference_audios is required.


Specifications

ItemDetail
OutputMP4 with native stereo audio
Frame rate24 fps
Resolution480p or 768p
Duration315 seconds (snaps to the model’s frame grid, so a 5s request lands at ~5.2s)
Reference imagesUp to 9
Reference videosUp to 3, 480p output only, 15s total budget
Reference audioUp to 3, trimmed to 15s each
SeedSupported

Pricing

Output video is billed per generated second, plus per-reference charges:

ItemPrice
Output video (480p)$0.06 / second
Output video (768p)$0.15 / second
Reference image$0.02 each
Reference audio$0.02 each
Reference video (480p)$0.06 / second

Explore the MiniMax H3 Open Weights Family

Authentication

For authentication details, please refer to the Authentication Guide.

API Endpoints

Submit Task & Query Result

set -euo pipefail

export WAVESPEED_API_KEY="your-api-key"

REQUEST_BODY=$(cat <<'JSON'
{
  "prompt": "A cinematic ocean wave at sunrise, highly detailed",
  "aspect_ratio": "16:9",
  "resolution": "480p",
  "duration": 5,
  "seed": -1
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/wavespeed-ai/minimax-h3/reference-to-video-lora" \
  -H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
  -H "Content-Type: application/json" \
  -d "${REQUEST_BODY}")

TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL=$(printf '%s' "${TASK}" | jq -r '.urls.get // empty')
if [ -z "${RESULT_URL}" ]; then RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"; fi

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body \
    "${RESULT_URL}" \
    -H "Authorization: Bearer ${WAVESPEED_API_KEY}")
  RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
  STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')

  case "${STATUS}" in
    completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
    failed|cancelled|timeout) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
    created|processing) sleep 2 ;;
    *) printf 'Unexpected status: %s
' "${STATUS}" >&2; exit 1 ;;
  esac
done

Parameters

Task Submission Parameters

Request Parameters

ParameterTypeRequiredDefaultRangeDescription
promptstringYes-Text description of the desired video. Refer to reference inputs as <Picture 1>..<Picture 9>, <Video 1>..<Video 3>, and <Audio 1>..<Audio 3>. Audio is generated natively together with the video.
reference_imagesarray<string>No-0 ~ 9 itemsReference image URLs (up to 9). At least one reference input is required.
reference_videosarray<string>No-0 ~ 3 itemsReference video URLs (up to 3, only supported at 480p resolution). Their soundtracks are used automatically. Total reference video duration is budgeted to 15 seconds; longer inputs are truncated fairly.
reference_audiosarray<string>No-0 ~ 3 itemsStandalone reference audio URLs (up to 3, trimmed to 15 seconds each).
aspect_ratiostringNo16:916:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21Output aspect ratio.
resolutionstringNo480p480p, 768pOutput video resolution. 768p is the model's native canvas; 480p is a faster, lower-cost tier.
durationintegerNo53, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15Output video duration in seconds.
seedintegerNo-1-The random seed to use for the generation. -1 means a random seed will be used.
lorasarray<object>No0 ~ 3 itemsList of LoRAs to apply (max 3).

Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
data.idstringUnique identifier for the prediction, Task Id
data.modelstringModel ID used for the prediction
data.outputsarrayOutput values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed)
data.urlsobjectObject containing related API endpoints
data.urls.getstringURL to retrieve the prediction result
data.statusstringStatus of the task: created, processing, completed, or failed
data.created_atstringISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”)
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds

Result Request Parameters

ParameterTypeRequiredDefaultDescription
idstringYes-Task ID

Result Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
dataobjectThe prediction data object containing all details
data.idstringUnique identifier for the prediction
data.modelstringModel ID used for the prediction
data.outputsarray<string | object>Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model.
data.urlsobjectObject containing related API endpoints
data.urls.getstringURL to poll for the prediction result
data.statusstringStatus: created, processing, completed, or failed
data.created_atstringISO timestamp of when the request was created
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds
© 2026 WaveSpeedAI. All rights reserved.