ElevenLabs Scribe V2 Speech-to-Text transcribes audio with automatic language detection, speaker labels, and word-level timestamps for transcription, subtitles, captions, meetings, and audio processing workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Idle
{
"text": "If the red of the second bow falls upon the green of the first, the result is to give a bow with an abnormally wide yellow band, since red and green light when mixed form yellow.",
"words": [
{
"end": 0.72,
"text": "If",
"type": "word",
"start": 0.64,
"speaker_id": null
},
{
"end": 0.78,
"text": " ",
"type": "spacing",
"start": 0.72,
"speaker_id": null
},
{
"end": 0.84,
"text": "the",
"type": "word",
"start": 0.78,
"speaker_id": null
},
{
"end": 0.9,
"text": " ",
"type": "spacing",
"start": 0.84,
"speaker_id": null
},
{
"end": 1.04,
"text": "red",
"type": "word",
"start": 0.9,
"speaker_id": null
},
{
"end": 1.08,
"text": " ",
"type": "spacing",
"start": 1.04,
"speaker_id": null
},
{
"end": 1.12,
"text": "of",
"type": "word",
"start": 1.08,
"speaker_id": null
},
{
"end": 1.16,
"text": " ",
"type": "spacing",
"start": 1.12,
"speaker_id": null
},
{
"end": 1.24,
"text": "the",
"type": "word",
"start": 1.16,
"speaker_id": null
},
{
"end": 1.28,
"text": " ",
"type": "spacing",
"start": 1.24,
"speaker_id": null
},
{
"end": 1.6,
"text": "second",
"type": "word",
"start": 1.28,
"speaker_id": null
},
{
"end": 1.64,
"text": " ",
"type": "spacing",
"start": 1.6,
"speaker_id": null
},
{
"end": 1.86,
"text": "bow",
"type": "word",
"start": 1.64,
"speaker_id": null
},
{
"end": 1.98,
"text": " ",
"type": "spacing",
"start": 1.86,
"speaker_id": null
},
{
"end": 2.14,
"text": "falls",
"type": "word",
"start": 1.98,
"speaker_id": null
},
{
"end": 2.18,
"text": " ",
"type": "spacing",
"start": 2.14,
"speaker_id": null
},
{
"end": 2.38,
"text": "upon",
"type": "word",
"start": 2.18,
"speaker_id": null
},
{
"end": 2.4,
"text": " ",
"type": "spacing",
"start": 2.38,
"speaker_id": null
},
{
"end": 2.48,
"text": "the",
"type": "word",
"start": 2.4,
"speaker_id": null
},
{
"end": 2.52,
"text": " ",
"type": "spacing",
"start": 2.48,
"speaker_id": null
},
{
"end": 2.7,
"text": "green",
"type": "word",
"start": 2.52,
"speaker_id": null
},
{
"end": 2.74,
"text": " ",
"type": "spacing",
"start": 2.7,
"speaker_id": null
},
{
"end": 2.8,
"text": "of",
"type": "word",
"start": 2.74,
"speaker_id": null
},
{
"end": 2.84,
"text": " ",
"type": "spacing",
"start": 2.8,
"speaker_id": null
},
{
"end": 2.92,
"text": "the",
"type": "word",
"start": 2.84,
"speaker_id": null
},
{
"end": 2.98,
"text": " ",
"type": "spacing",
"start": 2.92,
"speaker_id": null
},
{
"end": 3.38,
"text": "first,",
"type": "word",
"start": 2.98,
"speaker_id": null
},
{
"end": 3.38,
"text": " ",
"type": "spacing",
"start": 3.38,
"speaker_id": null
},
{
"end": 3.88,
"text": "the",
"type": "word",
"start": 3.82,
"speaker_id": null
},
{
"end": 3.94,
"text": " ",
"type": "spacing",
"start": 3.88,
"speaker_id": null
},
{
"end": 4.48,
"text": "result",
"type": "word",
"start": 3.94,
"speaker_id": null
},
{
"end": 4.56,
"text": " ",
"type": "spacing",
"start": 4.48,
"speaker_id": null
},
{
"end": 4.64,
"text": "is",
"type": "word",
"start": 4.56,
"speaker_id": null
},
{
"end": 4.7,
"text": " ",
"type": "spacing",
"start": 4.64,
"speaker_id": null
},
{
"end": 4.76,
"text": "to",
"type": "word",
"start": 4.7,
"speaker_id": null
},
{
"end": 4.8,
"text": " ",
"type": "spacing",
"start": 4.76,
"speaker_id": null
},
{
"end": 4.9,
"text": "give",
"type": "word",
"start": 4.8,
"speaker_id": null
},
{
"end": 4.94,
"text": " ",
"type": "spacing",
"start": 4.9,
"speaker_id": null
},
{
"end": 5.02,
"text": "a",
"type": "word",
"start": 4.94,
"speaker_id": null
},
{
"end": 5.06,
"text": " ",
"type": "spacing",
"start": 5.02,
"speaker_id": null
},
{
"end": 5.28,
"text": "bow",
"type": "word",
"start": 5.06,
"speaker_id": null
},
{
"end": 5.32,
"text": " ",
"type": "spacing",
"start": 5.28,
"speaker_id": null
},
{
"end": 5.42,
"text": "with",
"type": "word",
"start": 5.32,
"speaker_id": null
},
{
"end": 5.48,
"text": " ",
"type": "spacing",
"start": 5.42,
"speaker_id": null
},
{
"end": 5.54,
"text": "an",
"type": "word",
"start": 5.48,
"speaker_id": null
},
{
"end": 5.62,
"text": " ",
"type": "spacing",
"start": 5.54,
"speaker_id": null
},
{
"end": 6.08,
"text": "abnormally",
"type": "word",
"start": 5.62,
"speaker_id": null
},
{
"end": 6.12,
"text": " ",
"type": "spacing",
"start": 6.08,
"speaker_id": null
},
{
"end": 6.32,
"text": "wide",
"type": "word",
"start": 6.12,
"speaker_id": null
},
{
"end": 6.36,
"text": " ",
"type": "spacing",
"start": 6.32,
"speaker_id": null
},
{
"end": 6.56,
"text": "yellow",
"type": "word",
"start": 6.36,
"speaker_id": null
},
{
"end": 6.6,
"text": " ",
"type": "spacing",
"start": 6.56,
"speaker_id": null
},
{
"end": 6.88,
"text": "band,",
"type": "word",
"start": 6.6,
"speaker_id": null
},
{
"end": 7.22,
"text": " ",
"type": "spacing",
"start": 6.88,
"speaker_id": null
},
{
"end": 7.4,
"text": "since",
"type": "word",
"start": 7.22,
"speaker_id": null
},
{
"end": 7.46,
"text": " ",
"type": "spacing",
"start": 7.4,
"speaker_id": null
},
{
"end": 7.6,
"text": "red",
"type": "word",
"start": 7.46,
"speaker_id": null
},
{
"end": 7.64,
"text": " ",
"type": "spacing",
"start": 7.6,
"speaker_id": null
},
{
"end": 7.74,
"text": "and",
"type": "word",
"start": 7.64,
"speaker_id": null
},
{
"end": 7.78,
"text": " ",
"type": "spacing",
"start": 7.74,
"speaker_id": null
},
{
"end": 8,
"text": "green",
"type": "word",
"start": 7.78,
"speaker_id": null
},
{
"end": 8.06,
"text": " ",
"type": "spacing",
"start": 8,
"speaker_id": null
},
{
"end": 8.32,
"text": "light",
"type": "word",
"start": 8.06,
"speaker_id": null
},
{
"end": 8.38,
"text": " ",
"type": "spacing",
"start": 8.32,
"speaker_id": null
},
{
"end": 8.52,
"text": "when",
"type": "word",
"start": 8.38,
"speaker_id": null
},
{
"end": 8.56,
"text": " ",
"type": "spacing",
"start": 8.52,
"speaker_id": null
},
{
"end": 8.86,
"text": "mixed",
"type": "word",
"start": 8.56,
"speaker_id": null
},
{
"end": 8.92,
"text": " ",
"type": "spacing",
"start": 8.86,
"speaker_id": null
},
{
"end": 9.1,
"text": "form",
"type": "word",
"start": 8.92,
"speaker_id": null
},
{
"end": 9.16,
"text": " ",
"type": "spacing",
"start": 9.1,
"speaker_id": null
},
{
"end": 9.48,
"text": "yellow.",
"type": "word",
"start": 9.16,
"speaker_id": null
}
],
"language_code": "eng",
"language_probability": 1
}$0.01per run·~100 / $1
ElevenLabs Scribe V2 transcribes audio into structured text with automatic language detection, optional speaker diarization, word-level timestamps, and audio-event tagging. It is designed for interviews, meetings, podcasts, voice recordings, and other speech-to-text workflows where timing and speaker information matter.
Use optional keyterms to provide recognition hints for names, technical terms, brands, or other vocabulary that may be difficult to recognize from audio alone.
Speech-to-text transcription
Convert recorded speech into structured transcript data.
Automatic language detection
Let the model detect the spoken language automatically, or provide a language code when known.
Speaker diarization
Enable diarize to distinguish different speakers in multi-speaker recordings.
Word-level timestamps
Receive timing information for individual words for subtitles, editing, alignment, and search workflows.
Audio-event tagging
Detect non-speech events such as laughter when tag_audio_events is enabled.
Keyterm guidance
Provide up to 100 recognition hints for names, brands, technical vocabulary, or domain-specific terms.
| Parameter | Required | Description |
|---|---|---|
| audio | Yes | Audio input provided by upload or public direct file URL. |
| language_code | No | ISO-639-1 or ISO-639-3 language code. Omit for automatic language detection. |
| diarize | No | Identify and label different speakers in the recording. Default: false. |
| tag_audio_events | No | Include detected audio events such as laughter in the transcription result. Default: true. |
| keyterms | No | Recognition hints for important vocabulary. Supports up to 100 items; each term must be under 50 characters and contain at most 5 words. Additional pricing applies when any keyterms are provided. |
language_code when the spoken language is known, or omit it for automatic detection.diarize when multiple speakers should be identified separately.tag_audio_events enabled when non-speech events should be included.Pricing is based on the processed input audio duration.
| Configuration | Price per Input Minute |
|---|---|
| Standard transcription | $0.01 |
| With one or more keyterms | $0.012 |
Audio duration is charged proportionally rather than rounded up to full minutes.
| Audio Duration | Standard | With Keyterms |
|---|---|---|
| 30s | $0.005 | $0.006 |
| 1 min | $0.01 | $0.012 |
| 10 min | $0.10 | $0.12 |
| 30 min | $0.30 | $0.36 |
| 60 min | $0.60 | $0.72 |
Only the first 60 minutes of input audio are processed and billed.
Providing one or more keyterms applies a 1.2× multiplier to the transcription price. The number of keyterms does not add separate per-item charges.
language_code, diarize, and tag_audio_events do not add separate charges.
keyterms to improve recognition of specialized terminology.diarize when multiple people speak in the same recording.language_code when you already know the language and want to avoid relying on automatic detection.keyterms for uncommon names, brands, acronyms, product names, or technical vocabulary.audio is required.60 minutes of input audio are processed and billed.keyterms supports up to 100 items.50 characters and contain no more than 5 words.Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/elevenlabs/scribe-v2 with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Scribe v2 below.
set -euo pipefail
: "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"
REQUEST_BODY=$(cat <<'JSON'
{
"audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-audio/t-rex-roar.mp3",
"diarize": false,
"tag_audio_events": true
}
JSON
)
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/elevenlabs/scribe-v2" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $WAVESPEED_API_KEY" \
-d "$REQUEST_BODY")
TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
-H "Authorization: Bearer $WAVESPEED_API_KEY")
RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
case "$STATUS" in
completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
failed|cancelled|timeout|deleted) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
*) sleep 2 ;;
esac
doneconst submitUrl = "https://api.wavespeed.ai/api/v3/elevenlabs/scribe-v2";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');
async function requestJson(url, options = {}) {
const response = await fetch(url, options);
if (!response.ok) throw new Error(await response.text());
return response.json();
}
// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
method: "POST",
headers: {
"Authorization": `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
"audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-audio/t-rex-roar.mp3",
"diarize": false,
"tag_audio_events": true
}),
});
const task = body.data ?? body;
if (!task.id) throw new Error("Submission response did not contain a prediction id");
const resultUrl = `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;
// 2. Poll until the prediction finishes.
while (true) {
const resultBody = await requestJson(resultUrl, {
headers: { "Authorization": `Bearer ${apiKey}` },
});
const result = resultBody.data ?? resultBody;
if (result.status === "completed") {
console.log(result.outputs);
break;
}
if (["failed", "cancelled", "timeout", "deleted"].includes(result.status)) throw new Error(JSON.stringify(result));
await new Promise(resolve => setTimeout(resolve, 2000));
}import json
import os
import time
from urllib.request import Request, urlopen
api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
"audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-audio/t-rex-roar.mp3",
"diarize": False,
"tag_audio_events": True,
"keyterms": []
}
def request_json(url, data=None):
request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
with urlopen(request) as response:
return json.load(response)
# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/elevenlabs/scribe-v2", json.dumps(payload).encode())
task = body.get("data", body)
if not task.get("id"):
raise RuntimeError("Submission response did not contain a prediction id")
result_url = f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"
# 2. Poll until the prediction finishes.
while True:
result_body = request_json(result_url)
result = result_body.get("data", result_body)
status = result.get("status")
if status == "completed":
print(result.get("outputs", []))
break
if status in {"failed", "cancelled", "timeout", "deleted"}:
raise RuntimeError(result)
time.sleep(2)Scribe v2 is a ElevenLabs model for AI inference, exposed as a REST API on WaveSpeedAI. ElevenLabs Scribe V2 Speech-to-Text transcribes audio with automatic language detection, speaker labels, and word-level timestamps for transcription, subtitles, captions, meetings, and audio processing workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. You can call it programmatically or try it from the playground above.
POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates Python, JavaScript, and cURL examples for submitting requests and polling results. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/elevenlabs/elevenlabs-scribe-v2.
Scribe v2 starts at $0.01 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.
Key inputs: `audio`, `diarize`, `keyterms`, `language_code`, `tag_audio_events`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/elevenlabs/elevenlabs-scribe-v2.
Sign up for a free WaveSpeedAI account to claim starter credits, copy your API key from /accesskey, then call the endpoint shown in the API tab of the playground. The playground also auto-generates a code sample in Python, JavaScript, or cURL for the parameters you've set.
Commercial usage rights depend on the model's license, set by its provider (ElevenLabs). Check the provider's applicable terms and WaveSpeedAI's Terms of Service before commercial use.