Elevenlabs Scribe V2 API Documentation
Playground
Try it on WaveSpeedAI!ElevenLabs Scribe V2 Speech-to-Text transcribes audio with automatic language detection, speaker labels, and word-level timestamps for transcription, subtitles, captions, meetings, and audio processing workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Features
ElevenLabs Scribe V2 transcribes audio into structured text with automatic language detection, optional speaker diarization, word-level timestamps, and audio-event tagging. It is designed for interviews, meetings, podcasts, voice recordings, and other speech-to-text workflows where timing and speaker information matter.
Use optional keyterms to provide recognition hints for names, technical terms, brands, or other vocabulary that may be difficult to recognize from audio alone.
Why Choose This?
-
Speech-to-text transcription
Convert recorded speech into structured transcript data. -
Automatic language detection
Let the model detect the spoken language automatically, or provide a language code when known. -
Speaker diarization
Enablediarizeto distinguish different speakers in multi-speaker recordings. -
Word-level timestamps
Receive timing information for individual words for subtitles, editing, alignment, and search workflows. -
Audio-event tagging
Detect non-speech events such as laughter whentag_audio_eventsis enabled. -
Keyterm guidance
Provide up to100recognition hints for names, brands, technical vocabulary, or domain-specific terms.
Parameters
| Parameter | Required | Description |
|---|---|---|
| audio | Yes | Audio input provided by upload or public direct file URL. |
| language_code | No | ISO-639-1 or ISO-639-3 language code. Omit for automatic language detection. |
| diarize | No | Identify and label different speakers in the recording. Default: false. |
| tag_audio_events | No | Include detected audio events such as laughter in the transcription result. Default: true. |
| keyterms | No | Recognition hints for important vocabulary. Supports up to 100 items; each term must be under 50 characters and contain at most 5 words. Additional pricing applies when any keyterms are provided. |
How to Use
- Provide the audio — Upload a recording or supply a direct public audio-file URL.
- Set language optional — Provide
language_codewhen the spoken language is known, or omit it for automatic detection. - Enable diarization optional — Turn on
diarizewhen multiple speakers should be identified separately. - Configure audio events optional — Keep
tag_audio_eventsenabled when non-speech events should be included. - Add keyterms optional — Supply names, brands, technical vocabulary, or other recognition hints when needed.
- Submit — Process the recording and retrieve the transcript result.
Pricing
Pricing is based on the processed input audio duration.
| Configuration | Price per Input Minute |
|---|---|
| Standard transcription | $0.01 |
| With one or more keyterms | $0.012 |
Audio duration is charged proportionally rather than rounded up to full minutes.
Example Costs
| Audio Duration | Standard | With Keyterms |
|---|---|---|
| 30s | $0.005 | $0.006 |
| 1 min | $0.01 | $0.012 |
| 10 min | $0.10 | $0.12 |
| 30 min | $0.30 | $0.36 |
| 60 min | $0.60 | $0.72 |
Only the first 60 minutes of input audio are processed and billed.
Providing one or more keyterms applies a 1.2× multiplier to the transcription price. The number of keyterms does not add separate per-item charges.
language_code, diarize, and tag_audio_events do not add separate charges.
Best Use Cases
- Interviews — Transcribe conversations with optional speaker identification.
- Meetings — Create searchable transcripts with speaker labels and word timing.
- Podcasts — Convert spoken episodes into text for publishing, indexing, or subtitle workflows.
- Subtitle preparation — Use word-level timestamps to build caption and subtitle pipelines.
- Technical recordings — Use
keytermsto improve recognition of specialized terminology. - Content search and indexing — Convert audio archives into structured text for downstream processing.
Pro Tips
- Enable
diarizewhen multiple people speak in the same recording. - Provide
language_codewhen you already know the language and want to avoid relying on automatic detection. - Use
keytermsfor uncommon names, brands, acronyms, product names, or technical vocabulary. - Keep keyterms short and specific instead of using full explanatory sentences.
- Use a clean recording with clear speech for stronger transcription results.
- Supply a direct audio-file URL rather than a webpage containing an embedded audio player.
Notes
audiois required.- Only the first
60minutes of input audio are processed and billed. keytermssupports up to100items.- Each keyterm must be under
50characters and contain no more than5words. - Keyterms provide recognition guidance but do not guarantee exact spelling.
- The result includes transcript text, detected language information, and word-level timing data.
- When enabled, diarization adds speaker labels to the transcription result.
- Public URLs must point directly to an accessible audio file.
Related Models
- ElevenLabs Forced Alignment — Align an existing transcript with audio and return detailed timing information.
Authentication
For authentication details, please refer to the Authentication Guide.
API Endpoints
Submit Task & Query Result
set -euo pipefail
export WAVESPEED_API_KEY="your-api-key"
REQUEST_BODY=$(cat <<'JSON'
{
"audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-audio/t-rex-roar.mp3",
"diarize": false,
"tag_audio_events": true,
"keyterms": []
}
JSON
)
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/elevenlabs/scribe-v2" \
-H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
-H "Content-Type: application/json" \
-d "${REQUEST_BODY}")
TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body \
"${RESULT_URL}" \
-H "Authorization: Bearer ${WAVESPEED_API_KEY}")
RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')
case "${STATUS}" in
completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
failed|cancelled|timeout|deleted) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
*) sleep 2 ;;
esac
doneParameters
Task Submission Parameters
Request Parameters
| Parameter | Type | Required | Default | Range | Description |
|---|---|---|---|---|---|
| audio | string | Yes | - | 1 ~ unlimited characters · pattern: ^https?:// | Public direct URL of the audio file. Upload a file or supply a URL. |
| language_code | string | No | - | pattern: ^[a-z]{2,3}$ | ISO-639-1 or ISO-639-3 language code. Omit for automatic detection. |
| diarize | boolean | No | false | - | Identify different speakers. |
| tag_audio_events | boolean | No | true | - | Include audio events such as laughter. |
| keyterms | array<string> | No | [] | 0 ~ 100 items · each item: 1 ~ 49 characters · pattern: ^(?!.*[<>{}\[\]\\])\S+(?:\s+\S+){0,4}$ | Up to 100 recognition hints, each under 50 characters and at most 5 words. Additional charge applies. |
Response Parameters
| Parameter | Type | Description |
|---|---|---|
| code | integer | HTTP status code (e.g., 200 for success) |
| message | string | Status message (e.g., “success”) |
| data.id | string | Unique identifier for the prediction, Task Id |
| data.model | string | Model ID used for the prediction |
| data.outputs | array | Output values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed) |
| data.urls | object | Object containing related API endpoints |
| data.status | string | Task status. completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses. |
| data.created_at | string | ISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”) |
| data.error | string | Error message (empty if no error occurred) |
| data.timings | object | Object containing timing details |
| data.timings.inference | integer | Inference time in milliseconds |
Result Request Parameters
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
| id | string | Yes | - | Task ID |
Result Response Parameters
| Parameter | Type | Description |
|---|---|---|
| code | integer | HTTP status code (e.g., 200 for success) |
| message | string | Status message (e.g., “success”) |
| data | object | The prediction data object containing all details |
| data.id | string | Unique identifier for the prediction |
| data.model | string | Model ID used for the prediction |
| data.outputs | array<string | object> | Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model. |
| data.urls | object | Object containing related API endpoints |
| data.status | string | Status: completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses |
| data.created_at | string | ISO timestamp of when the request was created |
| data.error | string | Error message (empty if no error occurred) |
| data.timings | object | Object containing timing details |
| data.timings.inference | integer | Inference time in milliseconds |