Elevenlabs Scribe V2 API Documentation

Elevenlabs Scribe V2 API Documentation

Playground

Try it on WaveSpeedAI!

ElevenLabs Scribe V2 Speech-to-Text transcribes audio with automatic language detection, speaker labels, and word-level timestamps for transcription, subtitles, captions, meetings, and audio processing workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Features

ElevenLabs Scribe V2 transcribes audio into structured text with automatic language detection, optional speaker diarization, word-level timestamps, and audio-event tagging. It is designed for interviews, meetings, podcasts, voice recordings, and other speech-to-text workflows where timing and speaker information matter.

Use optional keyterms to provide recognition hints for names, technical terms, brands, or other vocabulary that may be difficult to recognize from audio alone.


Why Choose This?

  • Speech-to-text transcription
    Convert recorded speech into structured transcript data.

  • Automatic language detection
    Let the model detect the spoken language automatically, or provide a language code when known.

  • Speaker diarization
    Enable diarize to distinguish different speakers in multi-speaker recordings.

  • Word-level timestamps
    Receive timing information for individual words for subtitles, editing, alignment, and search workflows.

  • Audio-event tagging
    Detect non-speech events such as laughter when tag_audio_events is enabled.

  • Keyterm guidance
    Provide up to 100 recognition hints for names, brands, technical vocabulary, or domain-specific terms.


Parameters

ParameterRequiredDescription
audioYesAudio input provided by upload or public direct file URL.
language_codeNoISO-639-1 or ISO-639-3 language code. Omit for automatic language detection.
diarizeNoIdentify and label different speakers in the recording. Default: false.
tag_audio_eventsNoInclude detected audio events such as laughter in the transcription result. Default: true.
keytermsNoRecognition hints for important vocabulary. Supports up to 100 items; each term must be under 50 characters and contain at most 5 words. Additional pricing applies when any keyterms are provided.

How to Use

  1. Provide the audio — Upload a recording or supply a direct public audio-file URL.
  2. Set language optional — Provide language_code when the spoken language is known, or omit it for automatic detection.
  3. Enable diarization optional — Turn on diarize when multiple speakers should be identified separately.
  4. Configure audio events optional — Keep tag_audio_events enabled when non-speech events should be included.
  5. Add keyterms optional — Supply names, brands, technical vocabulary, or other recognition hints when needed.
  6. Submit — Process the recording and retrieve the transcript result.

Pricing

Pricing is based on the processed input audio duration.

ConfigurationPrice per Input Minute
Standard transcription$0.01
With one or more keyterms$0.012

Audio duration is charged proportionally rather than rounded up to full minutes.

Example Costs

Audio DurationStandardWith Keyterms
30s$0.005$0.006
1 min$0.01$0.012
10 min$0.10$0.12
30 min$0.30$0.36
60 min$0.60$0.72

Only the first 60 minutes of input audio are processed and billed.

Providing one or more keyterms applies a 1.2× multiplier to the transcription price. The number of keyterms does not add separate per-item charges.

language_code, diarize, and tag_audio_events do not add separate charges.


Best Use Cases

  • Interviews — Transcribe conversations with optional speaker identification.
  • Meetings — Create searchable transcripts with speaker labels and word timing.
  • Podcasts — Convert spoken episodes into text for publishing, indexing, or subtitle workflows.
  • Subtitle preparation — Use word-level timestamps to build caption and subtitle pipelines.
  • Technical recordings — Use keyterms to improve recognition of specialized terminology.
  • Content search and indexing — Convert audio archives into structured text for downstream processing.

Pro Tips

  • Enable diarize when multiple people speak in the same recording.
  • Provide language_code when you already know the language and want to avoid relying on automatic detection.
  • Use keyterms for uncommon names, brands, acronyms, product names, or technical vocabulary.
  • Keep keyterms short and specific instead of using full explanatory sentences.
  • Use a clean recording with clear speech for stronger transcription results.
  • Supply a direct audio-file URL rather than a webpage containing an embedded audio player.

Notes

  • audio is required.
  • Only the first 60 minutes of input audio are processed and billed.
  • keyterms supports up to 100 items.
  • Each keyterm must be under 50 characters and contain no more than 5 words.
  • Keyterms provide recognition guidance but do not guarantee exact spelling.
  • The result includes transcript text, detected language information, and word-level timing data.
  • When enabled, diarization adds speaker labels to the transcription result.
  • Public URLs must point directly to an accessible audio file.

Authentication

For authentication details, please refer to the Authentication Guide.

API Endpoints

Submit Task & Query Result

set -euo pipefail

export WAVESPEED_API_KEY="your-api-key"

REQUEST_BODY=$(cat <<'JSON'
{
  "audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-audio/t-rex-roar.mp3",
  "diarize": false,
  "tag_audio_events": true,
  "keyterms": []
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/elevenlabs/scribe-v2" \
  -H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
  -H "Content-Type: application/json" \
  -d "${REQUEST_BODY}")

TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body \
    "${RESULT_URL}" \
    -H "Authorization: Bearer ${WAVESPEED_API_KEY}")
  RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
  STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')

  case "${STATUS}" in
    completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
    failed|cancelled|timeout|deleted) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
    *) sleep 2 ;;
  esac
done

Parameters

Task Submission Parameters

Request Parameters

ParameterTypeRequiredDefaultRangeDescription
audiostringYes-1 ~ unlimited characters · pattern: ^https?://Public direct URL of the audio file. Upload a file or supply a URL.
language_codestringNo-pattern: ^[a-z]{2,3}$ISO-639-1 or ISO-639-3 language code. Omit for automatic detection.
diarizebooleanNofalse-Identify different speakers.
tag_audio_eventsbooleanNotrue-Include audio events such as laughter.
keytermsarray<string>No[]0 ~ 100 items · each item: 1 ~ 49 characters · pattern: ^(?!.*[<>{}\[\]\\])\S+(?:\s+\S+){0,4}$Up to 100 recognition hints, each under 50 characters and at most 5 words. Additional charge applies.

Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
data.idstringUnique identifier for the prediction, Task Id
data.modelstringModel ID used for the prediction
data.outputsarrayOutput values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed)
data.urlsobjectObject containing related API endpoints
data.statusstringTask status. completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses.
data.created_atstringISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”)
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds

Result Request Parameters

ParameterTypeRequiredDefaultDescription
idstringYes-Task ID

Result Response Parameters

ParameterTypeDescription
codeintegerHTTP status code (e.g., 200 for success)
messagestringStatus message (e.g., “success”)
dataobjectThe prediction data object containing all details
data.idstringUnique identifier for the prediction
data.modelstringModel ID used for the prediction
data.outputsarray<string | object>Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model.
data.urlsobjectObject containing related API endpoints
data.statusstringStatus: completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses
data.created_atstringISO timestamp of when the request was created
data.errorstringError message (empty if no error occurred)
data.timingsobjectObject containing timing details
data.timings.inferenceintegerInference time in milliseconds
© 2026 WaveSpeedAI. All rights reserved.