Minimax H3 Text To Video LoRA API Documentation

Minimax H3 Text To Video LoRA API Documentation

Playground

Try it on WaveSpeedAI!

MiniMax H3 Open Weights Text to Video with custom LoRA support generates coherent videos from text prompts, with 480P / 768P output, native stereo audio, 5-15 second duration, flexible aspect ratios, and per-second billing. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Features

Run the open-weights edition of MiniMax H3 on WaveSpeedAI’s own GPU infrastructure. This endpoint is separate from the official minimax/h3 API: same model family, independently hosted, with its own 480p/768p resolutions and per-second pricing.

MiniMax H3 is an omni-modal video model that generates picture and native stereo audio in a single pass — dialogue, sound effects, and music are produced together during generation, never added as a separate dubbing step. Describe both the visuals and the sound in one prompt and the model delivers a finished MP4 with a matching soundtrack.


How to Write a Great Prompt

H3 responds to prompts written as a timeline with a schedule, not a one-line description. The strongest prompts layer these blocks:

  1. Style — visual language up front: era, palette, texture, mood (e.g. “warm 1970s film look, soft grain, amber tones”).
  2. Timeline — timestamped beats across the duration: “[0s-2s] wide establishing shot … [2s-4s] slow push-in to the face …”. A clip sliced into beats hits every beat on schedule; one moment stretched across the whole duration looks static.
  3. Camera — state the movement explicitly, including when it should not move: “slow dolly-in, no cuts, no handheld”.
  4. Audio — put every sound in an Audio: line with entrance cues: “Audio: soft piano from 0s, a cello joins at 4s, gentle room tone throughout.” Omit this and the model picks a soundtrack for you.
  5. On-screen text — spell out every readable word in quotes: “the title reads ‘SUNRISE’”, then add “do not misspell, do not add other text, no subtitles.” Text you only describe comes back as letter-shaped noise.

Tips

  • Verbs beat adjectives — write the action (“she turns and smiles”), not just the vibe.
  • Iterate at 5s (composition and sound resolve there), then deliver at your final length.
  • A short negative list (“no dissolves, no captions, no extra people”) is free and prevents default drift.

Parameters

ParameterRequiredDescription
promptYesScene, action, camera movement, and an Audio: line for the soundtrack.
aspect_ratioNo16:9, 9:16, 1:1, 4:3, 3:4, 21:9, or 9:21. Default: 16:9.
resolutionNo480p (faster, lower cost) or 768p (native canvas). Default: 480p.
durationNoOutput length in seconds, 315. Default: 5.
seedNoFixed seed for reproducible output.
lorasNoUp to 3 LoRA weights, each {path, scale}; path is a LoRA file URL.

Specifications

ItemDetail
OutputMP4 with native stereo audio
Frame rate24 fps
Resolution480p or 768p
Duration315 seconds (snaps to the model’s frame grid, so a 5s request lands at ~5.2s)
SeedSupported

Pricing

Billed per generated second:

ResolutionPrice per second5s15s
480p$0.05$0.25$0.75
768p$0.10$0.50$1.50

Explore the MiniMax H3 Open Weights Family

Authentication

For authentication details, please refer to the Authentication Guide.

API Endpoints

Submit Task & Query Result

set -euo pipefail

export WAVESPEED_API_KEY="your-api-key"

REQUEST_BODY=$(cat <<'JSON'
{
  "prompt": "A cinematic ocean wave at sunrise, highly detailed",
  "aspect_ratio": "16:9",
  "resolution": "480p",
  "duration": 5,
  "seed": -1
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/wavespeed-ai/minimax-h3/text-to-video-lora" \
  -H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
  -H "Content-Type: application/json" \
  -d "${REQUEST_BODY}")

TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL=$(printf '%s' "${TASK}" | jq -r '.urls.get // empty')
if [ -z "${RESULT_URL}" ]; then RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"; fi

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body \
    "${RESULT_URL}" \
    -H "Authorization: Bearer ${WAVESPEED_API_KEY}")
  RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
  STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')

  case "${STATUS}" in
    completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
    failed|cancelled|timeout) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
    created|processing) sleep 2 ;;
    *) printf 'Unexpected status: %s
' "${STATUS}" >&2; exit 1 ;;
  esac
done

Parameters

Task Submission Parameters

Request Parameters

ParameterTypeRequiredDefaultRangeDescription
promptstringYes-Text description of the video scene, action, camera movement, and the desired soundtrack. Audio is generated natively together with the video.
aspect_ratiostringNo16:916:9, 9:16, 1:1, 4:3, 3:4, 21:9, 9:21Output aspect ratio.
resolutionstringNo480p480p, 768pOutput video resolution. 768p is the model's native canvas; 480p is a faster, lower-cost tier.
durationintegerNo53, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15Output video duration in seconds.
seedintegerNo-1-The random seed to use for the generation. -1 means a random seed will be used.
lorasarray<object>No0 ~ 3 itemsList of LoRAs to apply (max 3).

Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
data.idstringUnique identifier for the prediction, Task Id
data.modelstringModel ID used for the prediction
data.outputsarrayOutput values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed)
data.urlsobjectObject containing related API endpoints
data.urls.getstringURL to retrieve the prediction result
data.statusstringStatus of the task: created, processing, completed, or failed
data.created_atstringISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”)
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds

Result Request Parameters

ParameterTypeRequiredDefaultDescription
idstringYes-Task ID

Result Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
dataobjectThe prediction data object containing all details
data.idstringUnique identifier for the prediction
data.modelstringModel ID used for the prediction
data.outputsarray<string | object>Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model.
data.urlsobjectObject containing related API endpoints
data.urls.getstringURL to poll for the prediction result
data.statusstringStatus: created, processing, completed, or failed
data.created_atstringISO timestamp of when the request was created
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds
© 2026 WaveSpeedAI. All rights reserved.