MiniMax H3 Open Weights Video-Edit edits an input video from a natural-language prompt. The input video drives subject identity, composition, and motion while the model rewrites lighting, style, weather, environment, or specific elements as instructed, with native stereo audio generated in the same pass. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Idle
$0.125per run·~80 / $10
Transform the scene into a deep snowy winter night: dark blue evening sky, heavy snowfall with large visible snowflakes, thick snow covering the cobblestones, tables and awnings, glowing warm yellow light pouring from the cafe windows and the street lamps turned on, people wearing winter coats, scarves and beanies, breath visible in the cold air. Keep the same people, layout and camera motion.
MiniMax H3 Video Edit Open Weights edits an input video with prompt instructions while using the source video as the visual and motion foundation. You can guide the edit with optional reference images or reference audio, choose 480p or 768p, and decide whether to generate native audio or preserve the original input audio.
Prompt-guided video editing
Edit an existing video using natural-language instructions.
Source-video continuity
Use the input video to preserve subject identity, composition, and motion while applying visual changes.
Reference image support
Add reference images to guide subject identity, style, objects, or scene details.
Reference audio support
Add reference audio to guide native audio generation.
Native audio or preserved audio
Use generate_audio to generate native audio, or disable it to preserve the input video's audio track.
Flexible output settings
Choose output duration, aspect ratio, resolution, and seed for controlled generation.
| Parameter | Required | Description |
|---|---|---|
| prompt | Yes | Describe the edit you want applied to the input video. |
| video | Yes | URL of the input video to edit. It drives subject identity, composition, and motion while the prompt rewrites lighting, style, environment, or specific elements. |
| reference_images | No | Optional reference image URLs to guide the edit, such as subject identity, visual style, objects, or scene details. Supports up to 9 images. |
| reference_audios | No | Optional reference audio URLs to guide audio generation. Refer to them in the prompt as <Audio 1>..<Audio 3>. Supports up to 3 audio references. |
| resolution | No | Output video resolution: 480p or 768p. 768p is the model's native canvas; 480p is a faster, lower-cost tier. Default: 480p. |
| aspect_ratio | No | Output aspect ratio: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, or 9:21. If not specified, the output adapts to the input video. |
| duration | No | Output video duration in seconds. Options: 3 to 15. If not specified, duration is auto-detected from the input video. |
| generate_audio | No | Whether to generate native audio for the edited output. When set to false, the input video's audio track is preserved on the output instead. Default: true. |
| seed | No | Random seed for generation. A negative value means a random seed will be used. |
reference_images when identity, object, style, or scene consistency matters.reference_audios when native audio generation should follow a specific sound direction.480p for faster, lower-cost editing or 768p for the native canvas.3 to 15 seconds, or leave it empty to follow the input video duration.generate_audio enabled for native audio generation, or disable it to preserve the input video's audio track.Pricing is based on counted video duration, selected resolution, and the number of reference images and reference audios.
Counted video duration is calculated as input video duration + output duration.
Input video duration is rounded up and capped at 15 seconds. Output duration is capped at 15 seconds. If duration is not specified, the output duration follows the processed input duration with a minimum of 3 seconds.
| Item | Price |
|---|---|
| Video editing, 480p | $0.05 / counted second |
| Video editing, 768p | $0.125 / counted second |
| Reference image | $0.02 each |
| Reference audio | $0.02 each |
| Configuration | Cost |
|---|---|
| 5s input + 5s output at 480p | $0.50 |
| 5s input + 5s output at 768p | $1.25 |
| 10s input + 10s output at 480p | $1.00 |
| 10s input + 10s output at 768p | $2.50 |
| 5s input + 5s output at 480p + 2 reference images | $0.54 |
| 5s input + 5s output at 768p + 2 reference images + 1 reference audio | $1.31 |
aspect_ratio, generate_audio, prompt, and seed do not add separate charges.
480p.768p when higher-resolution output is needed.generate_audio is enabled.generate_audio to false when you want to preserve the input video's original audio.aspect_ratio empty when the output should follow the input video's format.480p for quick testing and 768p for higher-quality output.seed when comparing prompt or reference changes.prompt and video are required.reference_images supports up to 9 images.reference_audios supports up to 3 audio references.duration supports values from 3 to 15 seconds.duration is not specified, the output duration is auto-detected from the input video.generate_audio is false, the input video's audio track is preserved.Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/wavespeed-ai/minimax-h3/video-edit with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Minimax H3 Video Edit below.
set -euo pipefail
: "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"
REQUEST_BODY=$(cat <<'JSON'
{
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
"resolution": "480p",
"aspect_ratio": "16:9",
"duration": 3,
"generate_audio": true
}
JSON
)
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/wavespeed-ai/minimax-h3/video-edit" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $WAVESPEED_API_KEY" \
-d "$REQUEST_BODY")
TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
-H "Authorization: Bearer $WAVESPEED_API_KEY")
RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
case "$STATUS" in
completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
failed|cancelled|timeout|deleted) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
*) sleep 2 ;;
esac
doneconst submitUrl = "https://api.wavespeed.ai/api/v3/wavespeed-ai/minimax-h3/video-edit";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');
async function requestJson(url, options = {}) {
const response = await fetch(url, options);
if (!response.ok) throw new Error(await response.text());
return response.json();
}
// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
method: "POST",
headers: {
"Authorization": `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
"resolution": "480p",
"aspect_ratio": "16:9",
"duration": 3,
"generate_audio": true
}),
});
const task = body.data ?? body;
if (!task.id) throw new Error("Submission response did not contain a prediction id");
const resultUrl = `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;
// 2. Poll until the prediction finishes.
while (true) {
const resultBody = await requestJson(resultUrl, {
headers: { "Authorization": `Bearer ${apiKey}` },
});
const result = resultBody.data ?? resultBody;
if (result.status === "completed") {
console.log(result.outputs);
break;
}
if (["failed", "cancelled", "timeout", "deleted"].includes(result.status)) throw new Error(JSON.stringify(result));
await new Promise(resolve => setTimeout(resolve, 2000));
}import json
import os
import time
from urllib.request import Request, urlopen
api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
"prompt": "A cinematic shot of a city at sunset, soft golden light",
"video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
"resolution": "480p",
"aspect_ratio": "16:9",
"duration": 3,
"generate_audio": True
}
def request_json(url, data=None):
request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
with urlopen(request) as response:
return json.load(response)
# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/wavespeed-ai/minimax-h3/video-edit", json.dumps(payload).encode())
task = body.get("data", body)
if not task.get("id"):
raise RuntimeError("Submission response did not contain a prediction id")
result_url = f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"
# 2. Poll until the prediction finishes.
while True:
result_body = request_json(result_url)
result = result_body.get("data", result_body)
status = result.get("status")
if status == "completed":
print(result.get("outputs", []))
break
if status in {"failed", "cancelled", "timeout", "deleted"}:
raise RuntimeError(result)
time.sleep(2)Minimax H3 Video Edit is a WaveSpeedAI model for video editing, exposed as a REST API on WaveSpeedAI. MiniMax H3 Open Weights Video-Edit edits an input video from a natural-language prompt. The input video drives subject identity, composition, and motion while the model rewrites lighting, style, weather, environment, or specific elements as instructed, with native stereo audio generated in the same pass. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. You can call it programmatically or try it from the playground above.
POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates production-oriented Python, JavaScript, and cURL examples with timeouts, transient-error handling, and safe GET retries. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/wavespeed-ai/minimax-h3-video-edit.
Minimax H3 Video Edit starts at $0.25 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.
Key inputs: `prompt`, `video`, `aspect_ratio`, `resolution`, `duration`, `seed`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/wavespeed-ai/minimax-h3-video-edit.
Sign up for a free WaveSpeedAI account to claim starter credits, copy your API key from /accesskey, then call the endpoint shown in the API tab of the playground. The playground also auto-generates a code sample in Python, JavaScript, or cURL for the parameters you've set.
Commercial usage rights depend on the model's license, set by its provider (WaveSpeedAI). The license summary appears on the model card above; see WaveSpeedAI's Terms of Service for platform-level conditions.