Nano Banana 2.1 is LIVE — Google's Latest | Try Now →

elevenlabs/scribe-v2

ElevenLabs Scribe V2 Speech-to-Text transcribes audio with automatic language detection, speaker labels, and word-level timestamps for transcription, subtitles, captions, meetings, and audio processing workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Speech to Text
Input
Enable Safety Checker

Idle

{
  "text": "If the red of the second bow falls upon the green of the first, the result is to give a bow with an abnormally wide yellow band, since red and green light when mixed form yellow.",
  "words": [
    {
      "end": 0.72,
      "text": "If",
      "type": "word",
      "start": 0.64,
      "speaker_id": null
    },
    {
      "end": 0.78,
      "text": " ",
      "type": "spacing",
      "start": 0.72,
      "speaker_id": null
    },
    {
      "end": 0.84,
      "text": "the",
      "type": "word",
      "start": 0.78,
      "speaker_id": null
    },
    {
      "end": 0.9,
      "text": " ",
      "type": "spacing",
      "start": 0.84,
      "speaker_id": null
    },
    {
      "end": 1.04,
      "text": "red",
      "type": "word",
      "start": 0.9,
      "speaker_id": null
    },
    {
      "end": 1.08,
      "text": " ",
      "type": "spacing",
      "start": 1.04,
      "speaker_id": null
    },
    {
      "end": 1.12,
      "text": "of",
      "type": "word",
      "start": 1.08,
      "speaker_id": null
    },
    {
      "end": 1.16,
      "text": " ",
      "type": "spacing",
      "start": 1.12,
      "speaker_id": null
    },
    {
      "end": 1.24,
      "text": "the",
      "type": "word",
      "start": 1.16,
      "speaker_id": null
    },
    {
      "end": 1.28,
      "text": " ",
      "type": "spacing",
      "start": 1.24,
      "speaker_id": null
    },
    {
      "end": 1.6,
      "text": "second",
      "type": "word",
      "start": 1.28,
      "speaker_id": null
    },
    {
      "end": 1.64,
      "text": " ",
      "type": "spacing",
      "start": 1.6,
      "speaker_id": null
    },
    {
      "end": 1.86,
      "text": "bow",
      "type": "word",
      "start": 1.64,
      "speaker_id": null
    },
    {
      "end": 1.98,
      "text": " ",
      "type": "spacing",
      "start": 1.86,
      "speaker_id": null
    },
    {
      "end": 2.14,
      "text": "falls",
      "type": "word",
      "start": 1.98,
      "speaker_id": null
    },
    {
      "end": 2.18,
      "text": " ",
      "type": "spacing",
      "start": 2.14,
      "speaker_id": null
    },
    {
      "end": 2.38,
      "text": "upon",
      "type": "word",
      "start": 2.18,
      "speaker_id": null
    },
    {
      "end": 2.4,
      "text": " ",
      "type": "spacing",
      "start": 2.38,
      "speaker_id": null
    },
    {
      "end": 2.48,
      "text": "the",
      "type": "word",
      "start": 2.4,
      "speaker_id": null
    },
    {
      "end": 2.52,
      "text": " ",
      "type": "spacing",
      "start": 2.48,
      "speaker_id": null
    },
    {
      "end": 2.7,
      "text": "green",
      "type": "word",
      "start": 2.52,
      "speaker_id": null
    },
    {
      "end": 2.74,
      "text": " ",
      "type": "spacing",
      "start": 2.7,
      "speaker_id": null
    },
    {
      "end": 2.8,
      "text": "of",
      "type": "word",
      "start": 2.74,
      "speaker_id": null
    },
    {
      "end": 2.84,
      "text": " ",
      "type": "spacing",
      "start": 2.8,
      "speaker_id": null
    },
    {
      "end": 2.92,
      "text": "the",
      "type": "word",
      "start": 2.84,
      "speaker_id": null
    },
    {
      "end": 2.98,
      "text": " ",
      "type": "spacing",
      "start": 2.92,
      "speaker_id": null
    },
    {
      "end": 3.38,
      "text": "first,",
      "type": "word",
      "start": 2.98,
      "speaker_id": null
    },
    {
      "end": 3.38,
      "text": " ",
      "type": "spacing",
      "start": 3.38,
      "speaker_id": null
    },
    {
      "end": 3.88,
      "text": "the",
      "type": "word",
      "start": 3.82,
      "speaker_id": null
    },
    {
      "end": 3.94,
      "text": " ",
      "type": "spacing",
      "start": 3.88,
      "speaker_id": null
    },
    {
      "end": 4.48,
      "text": "result",
      "type": "word",
      "start": 3.94,
      "speaker_id": null
    },
    {
      "end": 4.56,
      "text": " ",
      "type": "spacing",
      "start": 4.48,
      "speaker_id": null
    },
    {
      "end": 4.64,
      "text": "is",
      "type": "word",
      "start": 4.56,
      "speaker_id": null
    },
    {
      "end": 4.7,
      "text": " ",
      "type": "spacing",
      "start": 4.64,
      "speaker_id": null
    },
    {
      "end": 4.76,
      "text": "to",
      "type": "word",
      "start": 4.7,
      "speaker_id": null
    },
    {
      "end": 4.8,
      "text": " ",
      "type": "spacing",
      "start": 4.76,
      "speaker_id": null
    },
    {
      "end": 4.9,
      "text": "give",
      "type": "word",
      "start": 4.8,
      "speaker_id": null
    },
    {
      "end": 4.94,
      "text": " ",
      "type": "spacing",
      "start": 4.9,
      "speaker_id": null
    },
    {
      "end": 5.02,
      "text": "a",
      "type": "word",
      "start": 4.94,
      "speaker_id": null
    },
    {
      "end": 5.06,
      "text": " ",
      "type": "spacing",
      "start": 5.02,
      "speaker_id": null
    },
    {
      "end": 5.28,
      "text": "bow",
      "type": "word",
      "start": 5.06,
      "speaker_id": null
    },
    {
      "end": 5.32,
      "text": " ",
      "type": "spacing",
      "start": 5.28,
      "speaker_id": null
    },
    {
      "end": 5.42,
      "text": "with",
      "type": "word",
      "start": 5.32,
      "speaker_id": null
    },
    {
      "end": 5.48,
      "text": " ",
      "type": "spacing",
      "start": 5.42,
      "speaker_id": null
    },
    {
      "end": 5.54,
      "text": "an",
      "type": "word",
      "start": 5.48,
      "speaker_id": null
    },
    {
      "end": 5.62,
      "text": " ",
      "type": "spacing",
      "start": 5.54,
      "speaker_id": null
    },
    {
      "end": 6.08,
      "text": "abnormally",
      "type": "word",
      "start": 5.62,
      "speaker_id": null
    },
    {
      "end": 6.12,
      "text": " ",
      "type": "spacing",
      "start": 6.08,
      "speaker_id": null
    },
    {
      "end": 6.32,
      "text": "wide",
      "type": "word",
      "start": 6.12,
      "speaker_id": null
    },
    {
      "end": 6.36,
      "text": " ",
      "type": "spacing",
      "start": 6.32,
      "speaker_id": null
    },
    {
      "end": 6.56,
      "text": "yellow",
      "type": "word",
      "start": 6.36,
      "speaker_id": null
    },
    {
      "end": 6.6,
      "text": " ",
      "type": "spacing",
      "start": 6.56,
      "speaker_id": null
    },
    {
      "end": 6.88,
      "text": "band,",
      "type": "word",
      "start": 6.6,
      "speaker_id": null
    },
    {
      "end": 7.22,
      "text": " ",
      "type": "spacing",
      "start": 6.88,
      "speaker_id": null
    },
    {
      "end": 7.4,
      "text": "since",
      "type": "word",
      "start": 7.22,
      "speaker_id": null
    },
    {
      "end": 7.46,
      "text": " ",
      "type": "spacing",
      "start": 7.4,
      "speaker_id": null
    },
    {
      "end": 7.6,
      "text": "red",
      "type": "word",
      "start": 7.46,
      "speaker_id": null
    },
    {
      "end": 7.64,
      "text": " ",
      "type": "spacing",
      "start": 7.6,
      "speaker_id": null
    },
    {
      "end": 7.74,
      "text": "and",
      "type": "word",
      "start": 7.64,
      "speaker_id": null
    },
    {
      "end": 7.78,
      "text": " ",
      "type": "spacing",
      "start": 7.74,
      "speaker_id": null
    },
    {
      "end": 8,
      "text": "green",
      "type": "word",
      "start": 7.78,
      "speaker_id": null
    },
    {
      "end": 8.06,
      "text": " ",
      "type": "spacing",
      "start": 8,
      "speaker_id": null
    },
    {
      "end": 8.32,
      "text": "light",
      "type": "word",
      "start": 8.06,
      "speaker_id": null
    },
    {
      "end": 8.38,
      "text": " ",
      "type": "spacing",
      "start": 8.32,
      "speaker_id": null
    },
    {
      "end": 8.52,
      "text": "when",
      "type": "word",
      "start": 8.38,
      "speaker_id": null
    },
    {
      "end": 8.56,
      "text": " ",
      "type": "spacing",
      "start": 8.52,
      "speaker_id": null
    },
    {
      "end": 8.86,
      "text": "mixed",
      "type": "word",
      "start": 8.56,
      "speaker_id": null
    },
    {
      "end": 8.92,
      "text": " ",
      "type": "spacing",
      "start": 8.86,
      "speaker_id": null
    },
    {
      "end": 9.1,
      "text": "form",
      "type": "word",
      "start": 8.92,
      "speaker_id": null
    },
    {
      "end": 9.16,
      "text": " ",
      "type": "spacing",
      "start": 9.1,
      "speaker_id": null
    },
    {
      "end": 9.48,
      "text": "yellow.",
      "type": "word",
      "start": 9.16,
      "speaker_id": null
    }
  ],
  "language_code": "eng",
  "language_probability": 1
}

$0.01per run·~100 / $1

ExamplesView all

Related Models

README

ElevenLabs Scribe V2

ElevenLabs Scribe V2 transcribes audio into structured text with automatic language detection, optional speaker diarization, word-level timestamps, and audio-event tagging. It is designed for interviews, meetings, podcasts, voice recordings, and other speech-to-text workflows where timing and speaker information matter.

Use optional keyterms to provide recognition hints for names, technical terms, brands, or other vocabulary that may be difficult to recognize from audio alone.

Why Choose This?

  • Speech-to-text transcription
    Convert recorded speech into structured transcript data.

  • Automatic language detection
    Let the model detect the spoken language automatically, or provide a language code when known.

  • Speaker diarization
    Enable diarize to distinguish different speakers in multi-speaker recordings.

  • Word-level timestamps
    Receive timing information for individual words for subtitles, editing, alignment, and search workflows.

  • Audio-event tagging
    Detect non-speech events such as laughter when tag_audio_events is enabled.

  • Keyterm guidance
    Provide up to 100 recognition hints for names, brands, technical vocabulary, or domain-specific terms.

Parameters

ParameterRequiredDescription
audioYesAudio input provided by upload or public direct file URL.
language_codeNoISO-639-1 or ISO-639-3 language code. Omit for automatic language detection.
diarizeNoIdentify and label different speakers in the recording. Default: false.
tag_audio_eventsNoInclude detected audio events such as laughter in the transcription result. Default: true.
keytermsNoRecognition hints for important vocabulary. Supports up to 100 items; each term must be under 50 characters and contain at most 5 words. Additional pricing applies when any keyterms are provided.

How to Use

  1. Provide the audio — Upload a recording or supply a direct public audio-file URL.
  2. Set language optional — Provide language_code when the spoken language is known, or omit it for automatic detection.
  3. Enable diarization optional — Turn on diarize when multiple speakers should be identified separately.
  4. Configure audio events optional — Keep tag_audio_events enabled when non-speech events should be included.
  5. Add keyterms optional — Supply names, brands, technical vocabulary, or other recognition hints when needed.
  6. Submit — Process the recording and retrieve the transcript result.

Pricing

Pricing is based on the processed input audio duration.

ConfigurationPrice per Input Minute
Standard transcription$0.01
With one or more keyterms$0.012

Audio duration is charged proportionally rather than rounded up to full minutes.

Example Costs

Audio DurationStandardWith Keyterms
30s$0.005$0.006
1 min$0.01$0.012
10 min$0.10$0.12
30 min$0.30$0.36
60 min$0.60$0.72

Only the first 60 minutes of input audio are processed and billed.

Providing one or more keyterms applies a 1.2× multiplier to the transcription price. The number of keyterms does not add separate per-item charges.

language_code, diarize, and tag_audio_events do not add separate charges.

Best Use Cases

  • Interviews — Transcribe conversations with optional speaker identification.
  • Meetings — Create searchable transcripts with speaker labels and word timing.
  • Podcasts — Convert spoken episodes into text for publishing, indexing, or subtitle workflows.
  • Subtitle preparation — Use word-level timestamps to build caption and subtitle pipelines.
  • Technical recordings — Use keyterms to improve recognition of specialized terminology.
  • Content search and indexing — Convert audio archives into structured text for downstream processing.

Pro Tips

  • Enable diarize when multiple people speak in the same recording.
  • Provide language_code when you already know the language and want to avoid relying on automatic detection.
  • Use keyterms for uncommon names, brands, acronyms, product names, or technical vocabulary.
  • Keep keyterms short and specific instead of using full explanatory sentences.
  • Use a clean recording with clear speech for stronger transcription results.
  • Supply a direct audio-file URL rather than a webpage containing an embedded audio player.

Notes

  • audio is required.
  • Only the first 60 minutes of input audio are processed and billed.
  • keyterms supports up to 100 items.
  • Each keyterm must be under 50 characters and contain no more than 5 words.
  • Keyterms provide recognition guidance but do not guarantee exact spelling.
  • The result includes transcript text, detected language information, and word-level timing data.
  • When enabled, diarization adds speaker labels to the transcription result.
  • Public URLs must point directly to an accessible audio file.

Related Models

Note:This website uses AI models provided by third parties. Documentation prices are for reference and may be outdated. The Generate button shows an estimate; the final task charge prevails.

Scribe v2 API — Quick start

Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/elevenlabs/scribe-v2 with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Scribe v2 below.

HTTP example
set -euo pipefail

: "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"

REQUEST_BODY=$(cat <<'JSON'
{
    "audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-audio/t-rex-roar.mp3",
    "diarize": false,
    "tag_audio_events": true
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/elevenlabs/scribe-v2" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $WAVESPEED_API_KEY" \
  -d "$REQUEST_BODY")

TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
    -H "Authorization: Bearer $WAVESPEED_API_KEY")
  RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
  STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
  case "$STATUS" in
    completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
    failed|cancelled|timeout|deleted) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
    *) sleep 2 ;;
  esac
done
Node.js example
const submitUrl = "https://api.wavespeed.ai/api/v3/elevenlabs/scribe-v2";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');

async function requestJson(url, options = {}) {
  const response = await fetch(url, options);
  if (!response.ok) throw new Error(await response.text());
  return response.json();
}

// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${apiKey}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
        "audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-audio/t-rex-roar.mp3",
        "diarize": false,
        "tag_audio_events": true
}),
});
const task = body.data ?? body;
if (!task.id) throw new Error("Submission response did not contain a prediction id");
const resultUrl = `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;

// 2. Poll until the prediction finishes.
while (true) {
  const resultBody = await requestJson(resultUrl, {
    headers: { "Authorization": `Bearer ${apiKey}` },
  });
  const result = resultBody.data ?? resultBody;
  if (result.status === "completed") {
    console.log(result.outputs);
    break;
  }
  if (["failed", "cancelled", "timeout", "deleted"].includes(result.status)) throw new Error(JSON.stringify(result));
  await new Promise(resolve => setTimeout(resolve, 2000));
}
Python example
import json
import os
import time
from urllib.request import Request, urlopen

api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
    "audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-audio/t-rex-roar.mp3",
    "diarize": False,
    "tag_audio_events": True,
    "keyterms": []
}

def request_json(url, data=None):
    request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
    with urlopen(request) as response:
        return json.load(response)

# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/elevenlabs/scribe-v2", json.dumps(payload).encode())
task = body.get("data", body)
if not task.get("id"):
    raise RuntimeError("Submission response did not contain a prediction id")
result_url = f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"

# 2. Poll until the prediction finishes.
while True:
    result_body = request_json(result_url)
    result = result_body.get("data", result_body)
    status = result.get("status")
    if status == "completed":
        print(result.get("outputs", []))
        break
    if status in {"failed", "cancelled", "timeout", "deleted"}:
        raise RuntimeError(result)
    time.sleep(2)

Scribe v2 API — Frequently asked questions

What is the Scribe v2 API?

Scribe v2 is a ElevenLabs model for AI inference, exposed as a REST API on WaveSpeedAI. ElevenLabs Scribe V2 Speech-to-Text transcribes audio with automatic language detection, speaker labels, and word-level timestamps for transcription, subtitles, captions, meetings, and audio processing workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. You can call it programmatically or try it from the playground above.

How do I call the Scribe v2 API?

POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates Python, JavaScript, and cURL examples for submitting requests and polling results. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/elevenlabs/elevenlabs-scribe-v2.

How much does Scribe v2 cost per run?

Scribe v2 starts at $0.01 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.

What inputs does Scribe v2 accept?

Key inputs: `audio`, `diarize`, `keyterms`, `language_code`, `tag_audio_events`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/elevenlabs/elevenlabs-scribe-v2.

How do I get started with the Scribe v2 API?

Sign up for a free WaveSpeedAI account to claim starter credits, copy your API key from /accesskey, then call the endpoint shown in the API tab of the playground. The playground also auto-generates a code sample in Python, JavaScript, or cURL for the parameters you've set.

Can I use Scribe v2 outputs commercially?

Commercial usage rights depend on the model's license, set by its provider (ElevenLabs). Check the provider's applicable terms and WaveSpeedAI's Terms of Service before commercial use.

llms.txt — elevenlabs/scribe-v2 for AI agents and LLMs