AI Video Captioner adds animated, word-by-word captions to any talking video: 12 caption styles, AI keyword highlights, matching emoji, optional silence and filler removal, about 100 languages with translation, plus an SRT subtitle file. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Idle
$0.08per run·~12 / $1
AI Video Captioner turns any talking video into a ready-to-post clip with animated, word-by-word captions. Speech is transcribed with word-level timing in about 100 languages, the important words are highlighted automatically, matching emoji pop in, and pauses or filler words can be cut out — all in one request.
translate_to.dictionary, or supply the script in transcript.| Parameter | Required | Description |
|---|---|---|
| video | Yes | The video to caption (URL or upload). Up to 2 hours is processed. |
| template | No | Caption style. Default bold-pop. |
| language | No | Language spoken in the video. auto detects it. |
| translate_to | No | Write the captions in this language instead. none keeps the spoken language. |
| highlight_keywords | No | Colour the most important words. Default true. |
| emoji | No | Add animated emoji that match the speech. Default true. |
| remove_silence | No | Cut pauses between sentences. Default false. |
| remove_filler_words | No | Cut hesitations such as "um" and "uh". Default false. |
| aspect_ratio | No | original, 9:16, 16:9, 1:1, 4:5 and more. Default original. |
| font | No | Font family; default uses the template's font. |
| font_size | No | Text size relative to the template, 0.5 to 2.0. |
| text_color / highlight_color / keyword_color | No | Colours as #RRGGBB or #RRGGBBAA. Empty uses the template. |
| stroke / stroke_color | No | Outline around the letters: none, thin, medium, thick. |
| shadow | No | none, soft, hard or glow. |
| position | No | Vertical centre of the captions, 0–100 % from the top. Leave it out to use the template. |
| text_case | No | upper for capitals, original to keep the spoken casing. |
| words_per_screen | No | Most words shown at once, 1–12. Leave it out to use the template. |
| dictionary | No | Names and terms to spell exactly. |
| transcript | No | The script of what is said, as plain text. See Using a transcript below. |
If you have the script, pass it in transcript and the captions use your wording — spelling, names, numbers and punctuation — while the timing still comes from the audio.
translate_to.dictionary or transcript? Use dictionary for a few names and terms the recognizer might misspell; use transcript when you have the whole script.outputs lists the captioned MP4 first, followed by its SRT subtitle file.
| Video length | Price |
|---|---|
| Per started minute | $0.08 |
Billing rounds the video's length up to whole minutes, for at most 120 minutes.
language explicitly for heavy accents, music beds or very short clips.dictionary for brand and product names so they are always spelled right.one-word and headline suit fast-paced shorts; minimal and bar suit interviews and tutorials.remove_silence with 9:16 to turn a recorded talk into a short in one step.Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/wavespeed-ai/ai-video-editor/video-captioner with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Ai Video Editor Video Captioner below.
set -euo pipefail
: "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"
REQUEST_BODY=$(cat <<'JSON'
{
"video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
"template": "bold-pop",
"language": "auto",
"translate_to": "none",
"highlight_keywords": true,
"emoji": true,
"remove_silence": false,
"remove_filler_words": false,
"aspect_ratio": "original"
}
JSON
)
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/wavespeed-ai/ai-video-editor/video-captioner" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $WAVESPEED_API_KEY" \
-d "$REQUEST_BODY")
TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
-H "Authorization: Bearer $WAVESPEED_API_KEY")
RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
case "$STATUS" in
completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
failed|cancelled|timeout|deleted) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
*) sleep 2 ;;
esac
doneconst submitUrl = "https://api.wavespeed.ai/api/v3/wavespeed-ai/ai-video-editor/video-captioner";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');
async function requestJson(url, options = {}) {
const response = await fetch(url, options);
if (!response.ok) throw new Error(await response.text());
return response.json();
}
// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
method: "POST",
headers: {
"Authorization": `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
"video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
"template": "bold-pop",
"language": "auto",
"translate_to": "none",
"highlight_keywords": true,
"emoji": true,
"remove_silence": false,
"remove_filler_words": false,
"aspect_ratio": "original"
}),
});
const task = body.data ?? body;
if (!task.id) throw new Error("Submission response did not contain a prediction id");
const resultUrl = `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;
// 2. Poll until the prediction finishes.
while (true) {
const resultBody = await requestJson(resultUrl, {
headers: { "Authorization": `Bearer ${apiKey}` },
});
const result = resultBody.data ?? resultBody;
if (result.status === "completed") {
console.log(result.outputs);
break;
}
if (["failed", "cancelled", "timeout", "deleted"].includes(result.status)) throw new Error(JSON.stringify(result));
await new Promise(resolve => setTimeout(resolve, 2000));
}import json
import os
import time
from urllib.request import Request, urlopen
api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
"video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
"template": "bold-pop",
"language": "auto",
"translate_to": "none",
"highlight_keywords": True,
"emoji": True,
"remove_silence": False,
"remove_filler_words": False,
"aspect_ratio": "original"
}
def request_json(url, data=None):
request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
with urlopen(request) as response:
return json.load(response)
# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/wavespeed-ai/ai-video-editor/video-captioner", json.dumps(payload).encode())
task = body.get("data", body)
if not task.get("id"):
raise RuntimeError("Submission response did not contain a prediction id")
result_url = f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"
# 2. Poll until the prediction finishes.
while True:
result_body = request_json(result_url)
result = result_body.get("data", result_body)
status = result.get("status")
if status == "completed":
print(result.get("outputs", []))
break
if status in {"failed", "cancelled", "timeout", "deleted"}:
raise RuntimeError(result)
time.sleep(2)Ai Video Editor Video Captioner is a WaveSpeedAI model for video editing, exposed as a REST API on WaveSpeedAI. AI Video Captioner adds animated, word-by-word captions to any talking video: 12 caption styles, AI keyword highlights, matching emoji, optional silence and filler removal, about 100 languages with translation, plus an SRT subtitle file. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. You can call it programmatically or try it from the playground above.
POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates Python, JavaScript, and cURL examples for submitting requests and polling results. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/wavespeed-ai/ai-video-editor-video-captioner.
Ai Video Editor Video Captioner starts at $0.08 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.
Key inputs: `video`, `aspect_ratio`, `emoji`, `highlight_keywords`, `language`, `remove_filler_words`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/wavespeed-ai/ai-video-editor-video-captioner.
Sign up for a free WaveSpeedAI account to claim starter credits, copy your API key from /accesskey, then call the endpoint shown in the API tab of the playground. The playground also auto-generates a code sample in Python, JavaScript, or cURL for the parameters you've set.
Commercial usage rights depend on the model's license, set by its provider (WaveSpeedAI). Check the provider's applicable terms and WaveSpeedAI's Terms of Service before commercial use.