# wavespeed-ai/ai-video-editor/video-captioner

> AI Video Captioner adds animated, word-by-word captions to any talking video: 12 caption styles, AI keyword highlights, matching emoji, optional silence and filler removal, about 100 languages with translation, plus an SRT subtitle file. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

## Overview

- **Endpoint**: `https://api.wavespeed.ai/api/v3/wavespeed-ai/ai-video-editor/video-captioner`
- **Polling/result URL**: `https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result`
- **Model ID**: `wavespeed-ai/ai-video-editor/video-captioner`
- **Category**: video-to-video

## API Information

This model can be used via our HTTP API or more conveniently via our client libraries.
The API is asynchronous: submit a prediction, then poll its result URL until it completes.

### Input Schema

The API accepts the following input parameters:

- **`video`** (`string`, _required_):
  The video to caption. Up to 2 hours is processed.

- **`template`** (`string`, _optional_):
  Caption look and animation: bold-pop (bold outline, spoken word pops), box (spoken word on a colour block), karaoke (words fill as spoken), one-word (one big word at a time), word-by-word (words appear as spoken), minimal (clean subtitles), focus (unspoken words dimmed), neon (glow), comic, headline (condensed type), bar (text on a translucent bar), playful.
  - Default: `"bold-pop"`
  - Options: "bold-pop", "box", "karaoke", "one-word", "word-by-word", "minimal", "focus", "neon", "comic", "headline", "bar", "playful"

- **`language`** (`string`, _optional_):
  Language spoken in the video. "auto" detects it; set it when detection is unreliable, for example with heavy accents or background music.
  - Default: `"auto"`
  - Options: "auto", "en", "zh-CN", "zh-TW", "yue", "es", "fr", "de", "it", "pt", "ru", "ja", "ko", "ar", "hi", "tr", "vi", "th", "id", "ms", "nl", "pl", "uk", "sv", "fi", "da", "no", "nn", "cs", "sk", "ro", "hu", "el", "bg", "hr", "sr", "sl", "bs", "mk", "sq", "lt", "lv", "et", "he", "fa", "ur", "ps", "sd", "bn", "as", "pa", "gu", "mr", "ne", "sa", "ta", "te", "kn", "ml", "si", "my", "km", "lo", "bo", "ka", "hy", "az", "kk", "uz", "tg", "tk", "mn", "ba", "tt", "be", "is", "fo", "cy", "br", "eu", "gl", "ca", "oc", "lb", "la", "mt", "af", "sw", "so", "am", "ha", "yo", "ln", "sn", "mg", "tl", "jw", "su", "haw", "mi", "ht", "yi"

- **`translate_to`** (`string`, _optional_):
  Translate the captions into this language. none keeps the spoken language.
  - Default: `"none"`
  - Options: "none", "en", "zh-CN", "zh-TW", "yue", "es", "fr", "de", "it", "pt", "ru", "ja", "ko", "ar", "hi", "tr", "vi", "th", "id", "ms", "nl", "pl", "uk", "sv", "fi", "da", "no", "nn", "cs", "sk", "ro", "hu", "el", "bg", "hr", "sr", "sl", "bs", "mk", "sq", "lt", "lv", "et", "he", "fa", "ur", "ps", "sd", "bn", "as", "pa", "gu", "mr", "ne", "sa", "ta", "te", "kn", "ml", "si", "my", "km", "lo", "bo", "ka", "hy", "az", "kk", "uz", "tg", "tk", "mn", "ba", "tt", "be", "is", "fo", "cy", "br", "eu", "gl", "ca", "oc", "lb", "la", "mt", "af", "sw", "so", "am", "ha", "yo", "ln", "sn", "mg", "tl", "jw", "su", "haw", "mi", "ht", "yi"

- **`highlight_keywords`** (`boolean`, _optional_):
  Colour the most important words.
  - Default: `true`

- **`emoji`** (`boolean`, _optional_):
  Add animated emoji that match what is said.
  - Default: `true`

- **`remove_silence`** (`boolean`, _optional_):
  Cut pauses between sentences.
  - Default: `false`

- **`remove_filler_words`** (`boolean`, _optional_):
  Cut hesitations such as um and uh.
  - Default: `false`

- **`aspect_ratio`** (`string`, _optional_):
  Output width:height. original keeps the source; other ratios keep the whole frame on a blurred fill.
  - Default: `"original"`
  - Options: "original", "9:16", "16:9", "1:1", "4:5", "5:4", "4:3", "3:4", "3:2", "2:3", "21:9"

- **`font`** (`string`, _optional_):
  Font family. default uses the template's font.
  - Default: `"default"`
  - Options: "default", "Montserrat-Black", "Montserrat-ExtraBold", "Poppins-Black", "Poppins-Bold", "Inter-ExtraBold", "Inter-SemiBold", "Anton-Regular", "BebasNeue-Regular", "Bangers-Regular", "LuckiestGuy-Regular", "ArchivoBlack-Regular", "NotoSans-Black"

- **`font_size`** (`number`, _optional_):
  Text size relative to the template, from 0.5 to 2.0.
  - Default: `1`
  - Range: `0.5` to `2`

- **`text_color`** (`string`, _optional_):
  Text colour as #RRGGBB or #RRGGBBAA. Empty uses the template.

- **`highlight_color`** (`string`, _optional_):
  Colour of the word being spoken (or of its box). Empty uses the template.

- **`keyword_color`** (`string`, _optional_):
  Colour of highlighted keywords. Empty uses the template.

- **`stroke`** (`string`, _optional_):
  Outline around the letters.
  - Default: `"default"`
  - Options: "default", "none", "thin", "medium", "thick"

- **`stroke_color`** (`string`, _optional_):
  Outline colour. Empty uses the template.

- **`shadow`** (`string`, _optional_):
  Text shadow.
  - Default: `"default"`
  - Options: "default", "none", "soft", "hard", "glow"

- **`position`** (`integer`, _optional_):
  Vertical centre of the captions as a percentage of the frame height from the top. Leave it empty to use the template's position.
  - Range: `0` to `100`

- **`text_case`** (`string`, _optional_):
  upper turns captions into capitals; original keeps the spoken casing.
  - Default: `"default"`
  - Options: "default", "upper", "original"

- **`words_per_screen`** (`integer`, _optional_):
  Most words shown at once, 1 to 12. Leave it empty to use the template.
  - Range: `1` to `12`

- **`dictionary`** (`array of string`, _optional_):
  Names and terms to spell exactly, such as brands or people.

- **`transcript`** (`string`, _optional_):
  The script of what is said, as plain text in the spoken language (no timestamps or speaker labels). Captions use its wording; timing comes from the audio. Ignored if most of it does not match the speech.



**Required Parameters Example**:

```json
{
  "video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4"
}
```

**Full Example**:

```json
{
  "video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4",
  "template": "bold-pop",
  "language": "auto",
  "translate_to": "none",
  "highlight_keywords": true,
  "emoji": true,
  "remove_silence": false,
  "remove_filler_words": false,
  "aspect_ratio": "original",
  "font": "default",
  "font_size": 1,
  "text_color": "A clear example input",
  "highlight_color": "example",
  "keyword_color": "example",
  "stroke": "default",
  "stroke_color": "example",
  "shadow": "default",
  "position": 0,
  "text_case": "default",
  "words_per_screen": 1,
  "dictionary": [],
  "transcript": "example"
}
```

### Result Data Schema

The `data` object returned by the API has the following fields:

- **`created_at`** (`string (date-time)`, _optional_):
  ISO timestamp of when the request was created (e.g., "2023-04-01T12:34:56.789Z").

- **`id`** (`string`, _optional_):
  Unique identifier for the prediction, the ID of the prediction to get.

- **`model`** (`string`, _optional_):
  Model ID used for the prediction.

- **`outputs`** (`array of string | object`, _optional_):
  Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model.

- **`status`** (`string`, _optional_):
  Status of the task: created, processing, completed, or failed.

- **`urls`** (`object`, _optional_):
  Object containing related API endpoints.



**Example `data` Object**:

```json
{
  "created_at": "example",
  "id": "example",
  "model": "example",
  "outputs": [],
  "status": "example",
  "urls": {}
}
```

## Usage Examples

The examples use `jq` to read JSON. Set your API key first:

```bash
set -euo pipefail
export WAVESPEED_API_KEY="your-api-key"
```

### 1. Submit a prediction

```bash
REQUEST_BODY=$(cat <<'JSON'
{
  "video": "https://interactive-examples.mdn.mozilla.net/media/cc0-videos/flower.mp4"
}
JSON
)

SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  --request POST \
  --url https://api.wavespeed.ai/api/v3/wavespeed-ai/ai-video-editor/video-captioner \
  --header "Authorization: Bearer ${WAVESPEED_API_KEY}" \
  --header "Content-Type: application/json" \
  --data "${REQUEST_BODY}")

printf '%s\n' "${SUBMIT_RESPONSE}" | jq .
```

The response contains the prediction ID in `data.id`.

### 2. Poll until complete and read `outputs`

```bash
PREDICTION_ID=$(printf '%s' "${SUBMIT_RESPONSE}" | jq -r '.data.id')
if [ -z "${PREDICTION_ID}" ] || [ "${PREDICTION_ID}" = "null" ]; then
  printf 'Submission response did not contain data.id\n' >&2
  exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"

while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body \
    --request GET \
    --url "${RESULT_URL}" \
    --header "Authorization: Bearer ${WAVESPEED_API_KEY}")

  RESULT=$(printf '%s' "${RESPONSE}" | jq -e '.data')
  STATUS=$(printf '%s' "${RESULT}" | jq -er '.status')
  case "${STATUS}" in
    completed)
      # Generated files are returned in the outputs array.
      printf '%s\n' "${RESULT}" | jq '.outputs'
      break
      ;;
    failed|cancelled|timeout|deleted)
      printf '%s\n' "${RESULT}" | jq '{status, error, code}'
      exit 1
      ;;
    *)
      sleep 2
      ;;
  esac
done
```

## Additional Resources

### Documentation

- [Model Playground](https://wavespeed.ai/models/wavespeed-ai/ai-video-editor/video-captioner)
- [API Documentation](https://wavespeed.ai/docs/docs-api/wavespeed-ai/ai-video-editor-video-captioner)
