GPT Image 2.5 is LIVE — Flare & Sunburst | Try in Image Generator →

elevenlabs/

ElevenLabs Forced Alignment aligns an existing transcript with an audio recording and returns character-level and word-level timestamps as structured JSON for subtitles, captions, dubbing, localization, and transcript synchronization workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

speech-to-text
Input
Enable Safety Checker

Idle

elevenlabs/forced-alignment preview unavailable

$0.3per run·~33 / $10

ExamplesView all

Hey everyone, Scribe V2 is now available on WaveSpeedAI, and this is a quick transcription test.

Related Models

README

ElevenLabs Forced Alignment

ElevenLabs Forced Alignment aligns an existing transcript with an audio recording. Provide an audio file and the transcript spoken in that audio, and the endpoint returns structured alignment data with character-level and word-level timestamps.

This endpoint does not generate a new audio file and does not create an automatic transcript. The prompt is the existing transcript to align.

Why Choose This?

  • Transcript alignment
    Align existing text with spoken audio.

  • Character-level timing
    Receive timestamp data for individual characters.

  • Word-level timing
    Retrieve word entries with start time, end time, and alignment loss.

  • Long audio support
    Supports audio inputs up to 10 hours, subject to validation.

  • Structured JSON output
    Use the returned alignment object for captions, subtitles, transcript viewers, or audio-text synchronization workflows.

Parameters

ParameterRequiredDescription
audioYesAudio recording URL or uploaded audio file. Supports up to 10 hours and 3 GB, subject to validation.
promptYesExisting transcript spoken in the audio. This is the text to align, not a generation instruction.

How to Use

  1. Upload audio — Provide the audio recording you want to align.
  2. Enter the transcript — Add the existing transcript spoken in the audio.
  3. Submit — Run the alignment request.
  4. Read the alignment JSON — Retrieve the structured alignment object from json.

Pricing

Pricing is based on input audio duration.

Billing is calculated per started hour. Audio duration is rounded up to the next whole hour, with a minimum charge of 1 hour and a maximum billed duration of 10 hours.

Billing UnitCost
Per started hour$0.30

Example Costs

Audio DurationBilled HoursCost
5 seconds1$0.30
30 minutes1$0.30
60 minutes1$0.30
61 minutes2$0.60
120 minutes2$0.60
10 hours10$3.00

Billing uses the measured duration of the input audio, not transcript length or the timestamp of the last aligned word.

Best Use Cases

  • Subtitle timing — Align transcript text with audio for subtitle or caption workflows.
  • Transcript viewers — Build word-level or character-level synchronized transcript interfaces.
  • Audio-text synchronization — Match spoken audio with existing text for playback highlighting.
  • Dataset preparation — Prepare timestamped transcript data for speech or media workflows.
  • Long-form audio alignment — Align interviews, podcasts, recordings, or narrated content.

Pro Tips

  • Use a transcript that closely matches the spoken audio.
  • Keep the transcript clean and ordered for better alignment.
  • Use the measured input audio duration to estimate billing.
  • Read alignment data from data.outputs[0].
  • Use word-level timestamps for subtitles and character-level timestamps for more precise highlighting.
  • Split very long audio when your workflow needs smaller alignment segments.

Notes

  • audio and prompt are required.
  • prompt must contain the existing transcript, not an instruction to generate text.
  • The endpoint returns structured alignment JSON only.
  • No voice ID is required.
  • The supported input limit is 10 hours.
  • Billing is rounded up to the next started hour and capped at 10 billed hours.
Note:This website uses AI models provided by third parties. Documentation prices are for reference and may be outdated. The Generate button shows an estimate; the final task charge prevails.

Forced Alignment API — Quick start

Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/elevenlabs/forced-alignment with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Forced Alignment below.

HTTP example
set -euo pipefail

: "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"

REQUEST_BODY=$(cat <<'JSON'
{
    "audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-audio/t-rex-roar.mp3",
    "prompt": "A cinematic shot of a city at sunset, soft golden light"
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/elevenlabs/forced-alignment" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $WAVESPEED_API_KEY" \
  -d "$REQUEST_BODY")

TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
    -H "Authorization: Bearer $WAVESPEED_API_KEY")
  RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
  STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
  case "$STATUS" in
    completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
    failed|cancelled|timeout|deleted) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
    *) sleep 2 ;;
  esac
done
Node.js example
const submitUrl = "https://api.wavespeed.ai/api/v3/elevenlabs/forced-alignment";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');

async function requestJson(url, options = {}) {
  const response = await fetch(url, options);
  if (!response.ok) throw new Error(await response.text());
  return response.json();
}

// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${apiKey}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
        "audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-audio/t-rex-roar.mp3",
        "prompt": "A cinematic shot of a city at sunset, soft golden light"
}),
});
const task = body.data ?? body;
if (!task.id) throw new Error("Submission response did not contain a prediction id");
const resultUrl = `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;

// 2. Poll until the prediction finishes.
while (true) {
  const resultBody = await requestJson(resultUrl, {
    headers: { "Authorization": `Bearer ${apiKey}` },
  });
  const result = resultBody.data ?? resultBody;
  if (result.status === "completed") {
    console.log(result.outputs);
    break;
  }
  if (["failed", "cancelled", "timeout", "deleted"].includes(result.status)) throw new Error(JSON.stringify(result));
  await new Promise(resolve => setTimeout(resolve, 2000));
}
Python example
import json
import os
import time
from urllib.request import Request, urlopen

api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
    "audio": "https://interactive-examples.mdn.mozilla.net/media/cc0-audio/t-rex-roar.mp3",
    "prompt": "A cinematic shot of a city at sunset, soft golden light"
}

def request_json(url, data=None):
    request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
    with urlopen(request) as response:
        return json.load(response)

# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/elevenlabs/forced-alignment", json.dumps(payload).encode())
task = body.get("data", body)
if not task.get("id"):
    raise RuntimeError("Submission response did not contain a prediction id")
result_url = f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"

# 2. Poll until the prediction finishes.
while True:
    result_body = request_json(result_url)
    result = result_body.get("data", result_body)
    status = result.get("status")
    if status == "completed":
        print(result.get("outputs", []))
        break
    if status in {"failed", "cancelled", "timeout", "deleted"}:
        raise RuntimeError(result)
    time.sleep(2)

Forced Alignment API — Frequently asked questions

What is the Forced Alignment API?

Forced Alignment is a ElevenLabs model for AI inference, exposed as a REST API on WaveSpeedAI. ElevenLabs Forced Alignment aligns an existing transcript with an audio recording and returns character-level and word-level timestamps as structured JSON for subtitles, captions, dubbing, localization, and transcript synchronization workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. You can call it programmatically or try it from the playground above.

How do I call the Forced Alignment API?

POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates production-oriented Python, JavaScript, and cURL examples with timeouts, transient-error handling, and safe GET retries. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/elevenlabs/elevenlabs-forced-alignment.

How much does Forced Alignment cost per run?

Forced Alignment starts at $0.30 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.

What inputs does Forced Alignment accept?

Key inputs: `prompt`, `audio`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/elevenlabs/elevenlabs-forced-alignment.

How do I get started with the Forced Alignment API?

Sign up for a free WaveSpeedAI account to claim starter credits, copy your API key from /accesskey, then call the endpoint shown in the API tab of the playground. The playground also auto-generates a code sample in Python, JavaScript, or cURL for the parameters you've set.

Can I use Forced Alignment outputs commercially?

Commercial usage rights depend on the model's license, set by its provider (ElevenLabs). The license summary appears on the model card above; see WaveSpeedAI's Terms of Service for platform-level conditions.

ElevenLabs Forced Alignment Audio Transcript Timestamps API on WaveSpeedAI