AI Video Editor Video Captioner API Documentation
Playground
Try it on WaveSpeedAI!AI Video Captioner adds animated, word-by-word captions to any talking video: 12 caption styles, AI keyword highlights, matching emoji, optional silence and filler removal, about 100 languages with translation, plus an SRT subtitle file. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Features
AI Video Captioner turns any talking video into a ready-to-post clip with animated, word-by-word captions. Speech is transcribed with word-level timing in about 100 languages, the important words are highlighted automatically, matching emoji pop in, and pauses or filler words can be cut out — all in one request.
Why It Stands Out
- 12 animated caption styles: bold-pop, box, karaoke, one-word, word-by-word, minimal, focus, neon, comic, headline, bar and playful.
- Word-level timing: each word lights up, pops or appears exactly when it is spoken.
- AI keyword highlights: the words that carry the point of each sentence get their own colour.
- Animated emoji: emoji that match what is said appear above the captions.
- Tighter edits: optionally cut pauses and hesitations such as “um” and “uh”.
- About 100 languages: auto-detected, including right-to-left and complex scripts such as Arabic, Hebrew, Hindi and Thai.
- Translation: caption the video in another language with
translate_to. - Exact spelling: add names and brand terms in
dictionary, or supply the script intranscript. - Any frame: keep the original frame or output 9:16, 1:1, 4:5 and more on a blurred fill.
- Subtitle file included: an SRT file comes with every video, and the subtitles are also embedded in the MP4 as a soft track.
Parameters
| Parameter | Required | Description |
|---|---|---|
| video | Yes | The video to caption (URL or upload). Up to 2 hours is processed. |
| template | No | Caption style. Default bold-pop. |
| language | No | Language spoken in the video. auto detects it. |
| translate_to | No | Write the captions in this language instead. none keeps the spoken language. |
| highlight_keywords | No | Colour the most important words. Default true. |
| emoji | No | Add animated emoji that match the speech. Default true. |
| remove_silence | No | Cut pauses between sentences. Default false. |
| remove_filler_words | No | Cut hesitations such as “um” and “uh”. Default false. |
| aspect_ratio | No | original, 9:16, 16:9, 1:1, 4:5 and more. Default original. |
| font | No | Font family; default uses the template’s font. |
| font_size | No | Text size relative to the template, 0.5 to 2.0. |
| text_color / highlight_color / keyword_color | No | Colours as #RRGGBB or #RRGGBBAA. Empty uses the template. |
| stroke / stroke_color | No | Outline around the letters: none, thin, medium, thick. |
| shadow | No | none, soft, hard or glow. |
| position | No | Vertical centre of the captions, 0–100 % from the top. Leave it out to use the template. |
| text_case | No | upper for capitals, original to keep the spoken casing. |
| words_per_screen | No | Most words shown at once, 1–12. Leave it out to use the template. |
| dictionary | No | Names and terms to spell exactly. |
| transcript | No | The script of what is said, as plain text. See Using a transcript below. |
Using a transcript
If you have the script, pass it in transcript and the captions use your wording — spelling, names, numbers and punctuation — while the timing still comes from the audio.
- Plain text only. No timestamps, numbering or speaker labels; line breaks and punctuation are fine, and punctuation helps the captions break at sentence ends.
- In the spoken language. Write what is said, as it is said. To caption in another language, keep the transcript in the spoken language and set
translate_to. - Match the video. Leave out lines that are not spoken (titles, stage directions). Small differences are fine, but if most of the script does not line up with the speech, it is ignored and the captions use the recognized words instead.
- Up to 20,000 characters.
dictionaryortranscript? Usedictionaryfor a few names and terms the recognizer might misspell; usetranscriptwhen you have the whole script.
Output
outputs lists the captioned MP4 first, followed by its SRT subtitle file.
How to Use
- Upload your video or paste a public URL.
- Pick a template and, if you like, adjust font, colours, position and words per screen.
- Choose extras — keyword highlights, emoji, silence and filler removal, translation.
- Click Run and download the captioned video and its subtitle file.
Pricing
| Video length | Price |
|---|---|
| Per started minute | $0.08 |
Billing rounds the video’s length up to whole minutes, for at most 120 minutes.
Examples
- 45-second clip → 1 minute → $0.08
- 3 min 20 s video → 4 minutes → $0.32
- 60-minute podcast → $4.80
Pro Tips
- Set
languageexplicitly for heavy accents, music beds or very short clips. - Use
dictionaryfor brand and product names so they are always spelled right. one-wordandheadlinesuit fast-paced shorts;minimalandbarsuit interviews and tutorials.- Combine
remove_silencewith9:16to turn a recorded talk into a short in one step.
Notes
- Videos longer than 2 hours are captioned for their first 2 hours.
- Ensure uploaded video URLs are publicly accessible.
Authentication
For authentication details, please refer to the Authentication Guide.
API Endpoints
Submit Task & Query Result
set -euo pipefail
export WAVESPEED_API_KEY="your-api-key"
REQUEST_BODY=$(cat <<'JSON'
{
"video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
"template": "bold-pop",
"language": "auto",
"translate_to": "none",
"highlight_keywords": true,
"emoji": true,
"remove_silence": false,
"remove_filler_words": false,
"aspect_ratio": "original"
}
JSON
)
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/wavespeed-ai/ai-video-editor/video-captioner" \
-H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
-H "Content-Type: application/json" \
-d "${REQUEST_BODY}")
TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body \
"${RESULT_URL}" \
-H "Authorization: Bearer ${WAVESPEED_API_KEY}")
RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')
case "${STATUS}" in
completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
failed|cancelled|timeout|deleted) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
*) sleep 2 ;;
esac
doneParameters
Task Submission Parameters
Request Parameters
| Parameter | Type | Required | Default | Range | Description |
|---|---|---|---|---|---|
| video | string | Yes | - | The video to caption. Up to 2 hours is processed. | |
| template | string | No | bold-pop | bold-pop, box, karaoke, one-word, word-by-word, minimal, focus, neon, comic, headline, bar, playful | Caption look and animation: bold-pop (bold outline, spoken word pops), box (spoken word on a colour block), karaoke (words fill as spoken), one-word (one big word at a time), word-by-word (words appear as spoken), minimal (clean subtitles), focus (unspoken words dimmed), neon (glow), comic, headline (condensed type), bar (text on a translucent bar), playful. |
| language | string | No | auto | auto, en, zh-CN, zh-TW, yue, es, fr, de, it, pt, ru, ja, ko, ar, hi, tr, vi, th, id, ms, nl, pl, uk, sv, fi, da, no, nn, cs, sk, ro, hu, el, bg, hr, sr, sl, bs, mk, sq, lt, lv, et, he, fa, ur, ps, sd, bn, as, pa, gu, mr, ne, sa, ta, te, kn, ml, si, my, km, lo, bo, ka, hy, az, kk, uz, tg, tk, mn, ba, tt, be, is, fo, cy, br, eu, gl, ca, oc, lb, la, mt, af, sw, so, am, ha, yo, ln, sn, mg, tl, jw, su, haw, mi, ht, yi | Language spoken in the video. "auto" detects it; set it when detection is unreliable, for example with heavy accents or background music. |
| translate_to | string | No | none | none, en, zh-CN, zh-TW, yue, es, fr, de, it, pt, ru, ja, ko, ar, hi, tr, vi, th, id, ms, nl, pl, uk, sv, fi, da, no, nn, cs, sk, ro, hu, el, bg, hr, sr, sl, bs, mk, sq, lt, lv, et, he, fa, ur, ps, sd, bn, as, pa, gu, mr, ne, sa, ta, te, kn, ml, si, my, km, lo, bo, ka, hy, az, kk, uz, tg, tk, mn, ba, tt, be, is, fo, cy, br, eu, gl, ca, oc, lb, la, mt, af, sw, so, am, ha, yo, ln, sn, mg, tl, jw, su, haw, mi, ht, yi | Translate the captions into this language. none keeps the spoken language. |
| highlight_keywords | boolean | No | true | - | Colour the most important words. |
| emoji | boolean | No | true | - | Add animated emoji that match what is said. |
| remove_silence | boolean | No | false | - | Cut pauses between sentences. |
| remove_filler_words | boolean | No | false | - | Cut hesitations such as um and uh. |
| aspect_ratio | string | No | original | original, 9:16, 16:9, 1:1, 4:5, 5:4, 4:3, 3:4, 3:2, 2:3, 21:9 | Output width:height. original keeps the source; other ratios keep the whole frame on a blurred fill. |
Response Parameters
| Parameter | Type | Description |
|---|---|---|
| code | integer | HTTP status code (e.g., 200 for success) |
| message | string | Status message (e.g., “success”) |
| data.id | string | Unique identifier for the prediction, Task Id |
| data.model | string | Model ID used for the prediction |
| data.outputs | array | Output values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed) |
| data.urls | object | Object containing related API endpoints |
| data.status | string | Task status. completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses. |
| data.created_at | string | ISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”) |
| data.error | string | Error message (empty if no error occurred) |
| data.timings | object | Object containing timing details |
| data.timings.inference | integer | Inference time in milliseconds |
Result Request Parameters
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
| id | string | Yes | - | Task ID |
Result Response Parameters
| Parameter | Type | Description |
|---|---|---|
| code | integer | HTTP status code (e.g., 200 for success) |
| message | string | Status message (e.g., “success”) |
| data | object | The prediction data object containing all details |
| data.id | string | Unique identifier for the prediction |
| data.model | string | Model ID used for the prediction |
| data.outputs | array<string | object> | Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model. |
| data.urls | object | Object containing related API endpoints |
| data.status | string | Status: completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses |
| data.created_at | string | ISO timestamp of when the request was created |
| data.error | string | Error message (empty if no error occurred) |
| data.timings | object | Object containing timing details |
| data.timings.inference | integer | Inference time in milliseconds |