Elevenlabs Dubbing API Documentation

Elevenlabs Dubbing API Documentation

Playground

Try it on WaveSpeedAI!

ElevenLabs Dubbing automatically translates and dubs video/audio content into different languages while preserving the original speakers’ voices. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Features

ElevenLabs Dubbing automatically translates and dubs video or audio content into different languages. Upload your content, select target language — the model transcribes, translates, and generates natural-sounding speech in the target language while preserving the original voice characteristics and timing.

REST inference API, best performance, no cold starts, affordable pricing.


Why Choose This?

  • Automatic translation and dubbing End-to-end pipeline: transcription → translation → voice synthesis in one step.

  • Voice preservation Maintains the original speaker’s voice characteristics in the dubbed output.

  • Multi-language support Supports English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese, Korean, and more.

  • Auto language detection Automatically detects the source language, or manually specify for better accuracy.

  • Video and audio support Works with both video files and audio-only content.


Parameters

ParameterRequiredDescription
videoNoSource video file to dub (upload or URL)
audioNoSource audio file to dub (upload or URL)
target_langYesTarget language for dubbing
source_langNoSource language (default: Auto for automatic detection)

Note: Either video or audio is required (upload one of them).

How to Use

  1. Upload your content — video or audio file (choose one).
  2. Select target language — the language you want the content dubbed into.
  3. Select source language (optional) — choose Auto for automatic detection, or specify manually.
  4. Run — submit and download the dubbed content.

Pricing

$0.01 per second of source media, rounded up to the next whole second.

DurationCost
30 seconds$0.30
1 minute$0.60
5 minutes$3.00
15 minutes$9.00

Billing Rules

  • Billed by the duration of the source media, prorated per second — no rounding up to whole minutes
  • If both video and audio are provided, only video is processed and billed
  • Maximum billed duration is 15 minutes; longer content is trimmed to the first 15 minutes
  • source_lang, target_lang and the other options do not affect pricing

Supported Languages

source_lang and target_lang take an iso639-1 or iso639-3 language code, not a language name — send es, not Spanish. source_lang also accepts auto (the default) to detect the source language automatically.

The two lists are not the same. A language you can dub from is not always a language you can dub into — Hebrew, Persian and Thai are supported sources but not targets, while Filipino is a supported target but not a source. Both lists below were verified against the live API.

Target languages (33)

CodeLanguageCodeLanguageCodeLanguage
arArabicdeGermanptPortuguese
bgBulgarianelGreekroRomanian
zhChinesehiHindiruRussian
hrCroatianhuHungarianskSlovak
csCzechidIndonesianesSpanish
daDanishitItaliansvSwedish
nlDutchjaJapanesetlTagalog
enEnglishkoKoreantaTamil
filFilipinomsMalaytrTurkish
fiFinnishnoNorwegianukUkrainian
frFrenchplPolishviVietnamese

Source languages (57, plus auto)

CodeLanguageCodeLanguageCodeLanguage
afAfrikaanselGreekfaPersian
arArabicheHebrewplPolish
hyArmenianhiHindiptPortuguese
azAzerbaijanihuHungarianroRomanian
beBelarusianisIcelandicruRussian
bsBosnianidIndonesiansrSerbian
bgBulgarianitItalianskSlovak
caCatalanjaJapaneseslSlovenian
zhChineseknKannadaesSpanish
hrCroatiankkKazakhswSwahili
csCzechkoKoreansvSwedish
daDanishlvLatviantlTagalog
nlDutchltLithuaniantaTamil
enEnglishmkMacedonianthThai
etEstonianmsMalaytrTurkish
fiFinnishmiMaoriukUkrainian
frFrenchmrMarathiurUrdu
glGalicianneNepaliviVietnamese
deGermannoNorwegiancyWelsh

Best Use Cases

  • Content Localization — Translate videos and podcasts for international audiences.
  • Marketing — Dub promotional content into multiple languages.
  • Education — Make educational content accessible in different languages.
  • Social Media — Reach global audiences with multilingual video content.
  • Corporate Communications — Translate internal videos and training materials.

Pro Tips

  • Use clean, high-quality source audio for best dubbing results.
  • Specify source language manually if auto-detection is inaccurate.
  • Shorter clips process faster — split long content for parallel processing.
  • Content with clear speech and minimal background noise produces better results.

Notes

  • Either video or audio is required — upload one of them.
  • If both video and audio are provided, the model will process video only.
  • Maximum duration is 15 minutes per job; longer content is automatically trimmed to the first 15 minutes.
  • For longer content, split into segments and process separately.

Authentication

For authentication details, please refer to the Authentication Guide.

API Endpoints

Submit Task & Query Result

set -euo pipefail

export WAVESPEED_API_KEY="your-api-key"

REQUEST_BODY=$(cat <<'JSON'
{
  "target_lang": "ar",
  "source_lang": "auto",
  "num_speakers": 0,
  "drop_background_audio": false
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/elevenlabs/dubbing" \
  -H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
  -H "Content-Type: application/json" \
  -d "${REQUEST_BODY}")

TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body \
    "${RESULT_URL}" \
    -H "Authorization: Bearer ${WAVESPEED_API_KEY}")
  RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
  STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')

  case "${STATUS}" in
    completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
    failed|cancelled|timeout|deleted) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
    *) sleep 2 ;;
  esac
done

Parameters

Task Submission Parameters

Request Parameters

ParameterTypeRequiredDefaultRangeDescription
target_langstringYes-ar, bg, zh, hr, cs, da, nl, en, fil, fi, fr, de, el, hi, hu, id, it, ja, ko, ms, no, pl, pt, ro, ru, sk, es, sv, tl, ta, tr, uk, viTarget language as an iso639-1 or iso639-3 code (e.g. es, fr, hi).
videostringNo-URL of the video file to dub. Either video or audio must be provided. If both are provided, video takes priority.
audiostringNo--URL of the audio file to dub. Either video or audio must be provided.
source_langstringNoautoauto, af, ar, hy, az, be, bs, bg, ca, zh, hr, cs, da, nl, en, et, fi, fr, gl, de, el, he, hi, hu, is, id, it, ja, kn, kk, ko, lv, lt, mk, ms, mi, mr, ne, no, fa, pl, pt, ro, ru, sr, sk, sl, es, sw, sv, tl, ta, th, tr, uk, ur, vi, cySource language as an iso639-1 or iso639-3 code (e.g. en, ja, zh). Use 'auto' to detect automatically.
num_speakersintegerNo00 ~ 32Number of speakers in the source. 0 detects automatically.
drop_background_audiobooleanNofalse-Drop background audio from the dubbed output.

Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
data.idstringUnique identifier for the prediction, Task Id
data.modelstringModel ID used for the prediction
data.outputsarrayOutput values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed)
data.urlsobjectObject containing related API endpoints
data.statusstringTask status. completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses.
data.created_atstringISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”)
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds

Result Request Parameters

ParameterTypeRequiredDefaultDescription
idstringYes-Task ID

Result Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
dataobjectThe prediction data object containing all details
data.idstringUnique identifier for the prediction
data.modelstringModel ID used for the prediction
data.outputsarray<string | object>Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model.
data.urlsobjectObject containing related API endpoints
data.statusstringStatus: completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses
data.created_atstringISO timestamp of when the request was created
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds
© 2026 WaveSpeedAI. All rights reserved.