Seedream 5.0 Flash is LIVE — Faster & Cheaper | Try Now →

wavespeed-ai/ai-video-editor/video-captioner

AI Video Captioner adds animated, word-by-word captions to any talking video: 12 caption styles, AI keyword highlights, matching emoji, optional silence and filler removal, about 100 languages with translation, plus an SRT subtitle file. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Video Editing
Input
Enable Safety Checker

Idle

Subtitles · SRT

    $0.08per run·~12 / $1

    Next:

    ExamplesView all

    Related Models

    README

    AI Video Captioner

    AI Video Captioner turns any talking video into a ready-to-post clip with animated, word-by-word captions. Speech is transcribed with word-level timing in about 100 languages, the important words are highlighted automatically, matching emoji pop in, and pauses or filler words can be cut out — all in one request.

    Why It Stands Out

    • 12 animated caption styles: bold-pop, box, karaoke, one-word, word-by-word, minimal, focus, neon, comic, headline, bar and playful.
    • Word-level timing: each word lights up, pops or appears exactly when it is spoken.
    • AI keyword highlights: the words that carry the point of each sentence get their own colour.
    • Animated emoji: emoji that match what is said appear above the captions.
    • Tighter edits: optionally cut pauses and hesitations such as "um" and "uh".
    • About 100 languages: auto-detected, including right-to-left and complex scripts such as Arabic, Hebrew, Hindi and Thai.
    • Translation: caption the video in another language with translate_to.
    • Exact spelling: add names and brand terms in dictionary, or supply the script in transcript.
    • Any frame: keep the original frame or output 9:16, 1:1, 4:5 and more on a blurred fill.
    • Subtitle file included: an SRT file comes with every video, and the subtitles are also embedded in the MP4 as a soft track.

    Parameters

    ParameterRequiredDescription
    videoYesThe video to caption (URL or upload). Up to 2 hours is processed.
    templateNoCaption style. Default bold-pop.
    languageNoLanguage spoken in the video. auto detects it.
    translate_toNoWrite the captions in this language instead. none keeps the spoken language.
    highlight_keywordsNoColour the most important words. Default true.
    emojiNoAdd animated emoji that match the speech. Default true.
    remove_silenceNoCut pauses between sentences. Default false.
    remove_filler_wordsNoCut hesitations such as "um" and "uh". Default false.
    aspect_ratioNooriginal, 9:16, 16:9, 1:1, 4:5 and more. Default original.
    fontNoFont family; default uses the template's font.
    font_sizeNoText size relative to the template, 0.5 to 2.0.
    text_color / highlight_color / keyword_colorNoColours as #RRGGBB or #RRGGBBAA. Empty uses the template.
    stroke / stroke_colorNoOutline around the letters: none, thin, medium, thick.
    shadowNonone, soft, hard or glow.
    positionNoVertical centre of the captions, 0–100 % from the top. Leave it out to use the template.
    text_caseNoupper for capitals, original to keep the spoken casing.
    words_per_screenNoMost words shown at once, 1–12. Leave it out to use the template.
    dictionaryNoNames and terms to spell exactly.
    transcriptNoThe script of what is said, as plain text. See Using a transcript below.

    Using a transcript

    If you have the script, pass it in transcript and the captions use your wording — spelling, names, numbers and punctuation — while the timing still comes from the audio.

    • Plain text only. No timestamps, numbering or speaker labels; line breaks and punctuation are fine, and punctuation helps the captions break at sentence ends.
    • In the spoken language. Write what is said, as it is said. To caption in another language, keep the transcript in the spoken language and set translate_to.
    • Match the video. Leave out lines that are not spoken (titles, stage directions). Small differences are fine, but if most of the script does not line up with the speech, it is ignored and the captions use the recognized words instead.
    • Part of the video is fine. A script of just the intro, say, sets the wording there; the rest of the speech is captioned from the recognized words.
    • Up to 20,000 characters.
    • dictionary or transcript? Use dictionary for a few names and terms the recognizer might misspell; use transcript when you have the whole script.

    Output

    outputs lists the captioned MP4 first, followed by its SRT subtitle file.

    How to Use

    1. Upload your video or paste a public URL.
    2. Pick a template and, if you like, adjust font, colours, position and words per screen.
    3. Choose extras — keyword highlights, emoji, silence and filler removal, translation.
    4. Click Run and download the captioned video and its subtitle file.

    Pricing

    Video lengthPrice
    Per started minute$0.08

    Billing rounds the video's length up to whole minutes, for at most 120 minutes.

    Examples

    • 45-second clip → 1 minute → $0.08
    • 3 min 20 s video → 4 minutes → $0.32
    • 60-minute podcast → $4.80

    Pro Tips

    • Set language explicitly for heavy accents, music beds or very short clips.
    • Use dictionary for brand and product names so they are always spelled right.
    • one-word and headline suit fast-paced shorts; minimal and bar suit interviews and tutorials.
    • Combine remove_silence with 9:16 to turn a recorded talk into a short in one step.

    Notes

    • Videos longer than 2 hours are captioned for their first 2 hours.
    • Ensure uploaded video URLs are publicly accessible.
    Note:This website uses AI models provided by third parties. Documentation prices are for reference and may be outdated. The Generate button shows an estimate; the final task charge prevails.

    Ai Video Editor Video Captioner API — Quick start

    Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/wavespeed-ai/ai-video-editor/video-captioner with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Ai Video Editor Video Captioner below.

    HTTP example
    set -euo pipefail
    
    : "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"
    
    REQUEST_BODY=$(cat <<'JSON'
    {
        "video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
        "template": "bold-pop",
        "language": "auto",
        "translate_to": "none",
        "highlight_keywords": true,
        "emoji": true,
        "remove_silence": false,
        "remove_filler_words": false,
        "aspect_ratio": "original"
    }
    JSON
    )
    
    # 1. Submit the prediction.
    SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
      -X POST "https://api.wavespeed.ai/api/v3/wavespeed-ai/ai-video-editor/video-captioner" \
      -H "Content-Type: application/json" \
      -H "Authorization: Bearer $WAVESPEED_API_KEY" \
      -d "$REQUEST_BODY")
    
    TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
    PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
    if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
      printf 'Submission response did not contain a prediction id
    ' >&2
      exit 1
    fi
    RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
    
    # 2. Poll until the prediction finishes.
    while true; do
      RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
        -H "Authorization: Bearer $WAVESPEED_API_KEY")
      RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
      STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
      case "$STATUS" in
        completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
        failed|cancelled|timeout|deleted) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
        *) sleep 2 ;;
      esac
    done
    Node.js example
    const submitUrl = "https://api.wavespeed.ai/api/v3/wavespeed-ai/ai-video-editor/video-captioner";
    const apiKey = process.env.WAVESPEED_API_KEY;
    if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');
    
    async function requestJson(url, options = {}) {
      const response = await fetch(url, options);
      if (!response.ok) throw new Error(await response.text());
      return response.json();
    }
    
    // 1. Submit the prediction.
    const body = await requestJson(submitUrl, {
      method: "POST",
      headers: {
        "Authorization": `Bearer ${apiKey}`,
        "Content-Type": "application/json",
      },
      body: JSON.stringify({
            "video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
            "template": "bold-pop",
            "language": "auto",
            "translate_to": "none",
            "highlight_keywords": true,
            "emoji": true,
            "remove_silence": false,
            "remove_filler_words": false,
            "aspect_ratio": "original"
    }),
    });
    const task = body.data ?? body;
    if (!task.id) throw new Error("Submission response did not contain a prediction id");
    const resultUrl = `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;
    
    // 2. Poll until the prediction finishes.
    while (true) {
      const resultBody = await requestJson(resultUrl, {
        headers: { "Authorization": `Bearer ${apiKey}` },
      });
      const result = resultBody.data ?? resultBody;
      if (result.status === "completed") {
        console.log(result.outputs);
        break;
      }
      if (["failed", "cancelled", "timeout", "deleted"].includes(result.status)) throw new Error(JSON.stringify(result));
      await new Promise(resolve => setTimeout(resolve, 2000));
    }
    Python example
    import json
    import os
    import time
    from urllib.request import Request, urlopen
    
    api_key = os.environ["WAVESPEED_API_KEY"]
    headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
    payload = {
        "video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
        "template": "bold-pop",
        "language": "auto",
        "translate_to": "none",
        "highlight_keywords": True,
        "emoji": True,
        "remove_silence": False,
        "remove_filler_words": False,
        "aspect_ratio": "original"
    }
    
    def request_json(url, data=None):
        request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
        with urlopen(request) as response:
            return json.load(response)
    
    # 1. Submit the prediction.
    body = request_json("https://api.wavespeed.ai/api/v3/wavespeed-ai/ai-video-editor/video-captioner", json.dumps(payload).encode())
    task = body.get("data", body)
    if not task.get("id"):
        raise RuntimeError("Submission response did not contain a prediction id")
    result_url = f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"
    
    # 2. Poll until the prediction finishes.
    while True:
        result_body = request_json(result_url)
        result = result_body.get("data", result_body)
        status = result.get("status")
        if status == "completed":
            print(result.get("outputs", []))
            break
        if status in {"failed", "cancelled", "timeout", "deleted"}:
            raise RuntimeError(result)
        time.sleep(2)

    Ai Video Editor Video Captioner API — Frequently asked questions

    What is the Ai Video Editor Video Captioner API?

    Ai Video Editor Video Captioner is a WaveSpeedAI model for video editing, exposed as a REST API on WaveSpeedAI. AI Video Captioner adds animated, word-by-word captions to any talking video: 12 caption styles, AI keyword highlights, matching emoji, optional silence and filler removal, about 100 languages with translation, plus an SRT subtitle file. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. You can call it programmatically or try it from the playground above.

    How do I call the Ai Video Editor Video Captioner API?

    POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates Python, JavaScript, and cURL examples for submitting requests and polling results. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/wavespeed-ai/ai-video-editor-video-captioner.

    How much does Ai Video Editor Video Captioner cost per run?

    Ai Video Editor Video Captioner starts at $0.08 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.

    What inputs does Ai Video Editor Video Captioner accept?

    Key inputs: `video`, `aspect_ratio`, `emoji`, `highlight_keywords`, `language`, `remove_filler_words`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/wavespeed-ai/ai-video-editor-video-captioner.

    How do I get started with the Ai Video Editor Video Captioner API?

    Sign up for a free WaveSpeedAI account to claim starter credits, copy your API key from /accesskey, then call the endpoint shown in the API tab of the playground. The playground also auto-generates a code sample in Python, JavaScript, or cURL for the parameters you've set.

    Can I use Ai Video Editor Video Captioner outputs commercially?

    Commercial usage rights depend on the model's license, set by its provider (WaveSpeedAI). Check the provider's applicable terms and WaveSpeedAI's Terms of Service before commercial use.

    llms.txt — wavespeed-ai/ai-video-editor/video-captioner for AI agents and LLMs