AI Video Editor Video Captioner API Documentation

AI Video Editor Video Captioner API Documentation

Playground

Try it on WaveSpeedAI!

AI Video Captioner adds animated, word-by-word captions to any talking video: 12 caption styles, AI keyword highlights, matching emoji, optional silence and filler removal, about 100 languages with translation, plus an SRT subtitle file. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Features

AI Video Captioner turns any talking video into a ready-to-post clip with animated, word-by-word captions. Speech is transcribed with word-level timing in about 100 languages, the important words are highlighted automatically, matching emoji pop in, and pauses or filler words can be cut out — all in one request.


Why It Stands Out

  • 12 animated caption styles: bold-pop, box, karaoke, one-word, word-by-word, minimal, focus, neon, comic, headline, bar and playful.
  • Word-level timing: each word lights up, pops or appears exactly when it is spoken.
  • AI keyword highlights: the words that carry the point of each sentence get their own colour.
  • Animated emoji: emoji that match what is said appear above the captions.
  • Tighter edits: optionally cut pauses and hesitations such as “um” and “uh”.
  • About 100 languages: auto-detected, including right-to-left and complex scripts such as Arabic, Hebrew, Hindi and Thai.
  • Translation: caption the video in another language with translate_to.
  • Exact spelling: add names and brand terms in dictionary, or supply the script in transcript.
  • Any frame: keep the original frame or output 9:16, 1:1, 4:5 and more on a blurred fill.
  • Subtitle file included: an SRT file comes with every video, and the subtitles are also embedded in the MP4 as a soft track.

Parameters

ParameterRequiredDescription
videoYesThe video to caption (URL or upload). Up to 2 hours is processed.
templateNoCaption style. Default bold-pop.
languageNoLanguage spoken in the video. auto detects it.
translate_toNoWrite the captions in this language instead. none keeps the spoken language.
highlight_keywordsNoColour the most important words. Default true.
emojiNoAdd animated emoji that match the speech. Default true.
remove_silenceNoCut pauses between sentences. Default false.
remove_filler_wordsNoCut hesitations such as “um” and “uh”. Default false.
aspect_ratioNooriginal, 9:16, 16:9, 1:1, 4:5 and more. Default original.
fontNoFont family; default uses the template’s font.
font_sizeNoText size relative to the template, 0.5 to 2.0.
text_color / highlight_color / keyword_colorNoColours as #RRGGBB or #RRGGBBAA. Empty uses the template.
stroke / stroke_colorNoOutline around the letters: none, thin, medium, thick.
shadowNonone, soft, hard or glow.
positionNoVertical centre of the captions, 0–100 % from the top. Leave it out to use the template.
text_caseNoupper for capitals, original to keep the spoken casing.
words_per_screenNoMost words shown at once, 1–12. Leave it out to use the template.
dictionaryNoNames and terms to spell exactly.
transcriptNoThe script of what is said, as plain text. See Using a transcript below.

Using a transcript

If you have the script, pass it in transcript and the captions use your wording — spelling, names, numbers and punctuation — while the timing still comes from the audio.

  • Plain text only. No timestamps, numbering or speaker labels; line breaks and punctuation are fine, and punctuation helps the captions break at sentence ends.
  • In the spoken language. Write what is said, as it is said. To caption in another language, keep the transcript in the spoken language and set translate_to.
  • Match the video. Leave out lines that are not spoken (titles, stage directions). Small differences are fine, but if most of the script does not line up with the speech, it is ignored and the captions use the recognized words instead.
  • Up to 20,000 characters.
  • dictionary or transcript? Use dictionary for a few names and terms the recognizer might misspell; use transcript when you have the whole script.

Output

outputs lists the captioned MP4 first, followed by its SRT subtitle file.


How to Use

  1. Upload your video or paste a public URL.
  2. Pick a template and, if you like, adjust font, colours, position and words per screen.
  3. Choose extras — keyword highlights, emoji, silence and filler removal, translation.
  4. Click Run and download the captioned video and its subtitle file.

Pricing

Video lengthPrice
Per started minute$0.08

Billing rounds the video’s length up to whole minutes, for at most 120 minutes.

Examples

  • 45-second clip → 1 minute → $0.08
  • 3 min 20 s video → 4 minutes → $0.32
  • 60-minute podcast → $4.80

Pro Tips

  • Set language explicitly for heavy accents, music beds or very short clips.
  • Use dictionary for brand and product names so they are always spelled right.
  • one-word and headline suit fast-paced shorts; minimal and bar suit interviews and tutorials.
  • Combine remove_silence with 9:16 to turn a recorded talk into a short in one step.

Notes

  • Videos longer than 2 hours are captioned for their first 2 hours.
  • Ensure uploaded video URLs are publicly accessible.

Authentication

For authentication details, please refer to the Authentication Guide.

API Endpoints

Submit Task & Query Result

set -euo pipefail

export WAVESPEED_API_KEY="your-api-key"

REQUEST_BODY=$(cat <<'JSON'
{
  "video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
  "template": "bold-pop",
  "language": "auto",
  "translate_to": "none",
  "highlight_keywords": true,
  "emoji": true,
  "remove_silence": false,
  "remove_filler_words": false,
  "aspect_ratio": "original"
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/wavespeed-ai/ai-video-editor/video-captioner" \
  -H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
  -H "Content-Type: application/json" \
  -d "${REQUEST_BODY}")

TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body \
    "${RESULT_URL}" \
    -H "Authorization: Bearer ${WAVESPEED_API_KEY}")
  RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
  STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')

  case "${STATUS}" in
    completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
    failed|cancelled|timeout|deleted) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
    *) sleep 2 ;;
  esac
done

Parameters

Task Submission Parameters

Request Parameters

ParameterTypeRequiredDefaultRangeDescription
videostringYes-The video to caption. Up to 2 hours is processed.
templatestringNobold-popbold-pop, box, karaoke, one-word, word-by-word, minimal, focus, neon, comic, headline, bar, playfulCaption look and animation: bold-pop (bold outline, spoken word pops), box (spoken word on a colour block), karaoke (words fill as spoken), one-word (one big word at a time), word-by-word (words appear as spoken), minimal (clean subtitles), focus (unspoken words dimmed), neon (glow), comic, headline (condensed type), bar (text on a translucent bar), playful.
languagestringNoautoauto, en, zh-CN, zh-TW, yue, es, fr, de, it, pt, ru, ja, ko, ar, hi, tr, vi, th, id, ms, nl, pl, uk, sv, fi, da, no, nn, cs, sk, ro, hu, el, bg, hr, sr, sl, bs, mk, sq, lt, lv, et, he, fa, ur, ps, sd, bn, as, pa, gu, mr, ne, sa, ta, te, kn, ml, si, my, km, lo, bo, ka, hy, az, kk, uz, tg, tk, mn, ba, tt, be, is, fo, cy, br, eu, gl, ca, oc, lb, la, mt, af, sw, so, am, ha, yo, ln, sn, mg, tl, jw, su, haw, mi, ht, yiLanguage spoken in the video. "auto" detects it; set it when detection is unreliable, for example with heavy accents or background music.
translate_tostringNononenone, en, zh-CN, zh-TW, yue, es, fr, de, it, pt, ru, ja, ko, ar, hi, tr, vi, th, id, ms, nl, pl, uk, sv, fi, da, no, nn, cs, sk, ro, hu, el, bg, hr, sr, sl, bs, mk, sq, lt, lv, et, he, fa, ur, ps, sd, bn, as, pa, gu, mr, ne, sa, ta, te, kn, ml, si, my, km, lo, bo, ka, hy, az, kk, uz, tg, tk, mn, ba, tt, be, is, fo, cy, br, eu, gl, ca, oc, lb, la, mt, af, sw, so, am, ha, yo, ln, sn, mg, tl, jw, su, haw, mi, ht, yiTranslate the captions into this language. none keeps the spoken language.
highlight_keywordsbooleanNotrue-Colour the most important words.
emojibooleanNotrue-Add animated emoji that match what is said.
remove_silencebooleanNofalse-Cut pauses between sentences.
remove_filler_wordsbooleanNofalse-Cut hesitations such as um and uh.
aspect_ratiostringNooriginaloriginal, 9:16, 16:9, 1:1, 4:5, 5:4, 4:3, 3:4, 3:2, 2:3, 21:9Output width:height. original keeps the source; other ratios keep the whole frame on a blurred fill.

Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
data.idstringUnique identifier for the prediction, Task Id
data.modelstringModel ID used for the prediction
data.outputsarrayOutput values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed)
data.urlsobjectObject containing related API endpoints
data.statusstringTask status. completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses.
data.created_atstringISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”)
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds

Result Request Parameters

ParameterTypeRequiredDefaultDescription
idstringYes-Task ID

Result Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
dataobjectThe prediction data object containing all details
data.idstringUnique identifier for the prediction
data.modelstringModel ID used for the prediction
data.outputsarray<string | object>Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model.
data.urlsobjectObject containing related API endpoints
data.statusstringStatus: completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses
data.created_atstringISO timestamp of when the request was created
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds
© 2026 WaveSpeedAI. All rights reserved.