Elevenlabs Eleven V4 API Documentation
Playground
Try it on WaveSpeedAI!Eleven V4 Text-to-Speech converts text into expressive multilingual speech with configurable voice selection, stability, and similarity controls for narration, dialogue, localization, virtual assistants, and voice production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Features
Eleven V4 converts text into expressive, natural-sounding speech for narration, voiceovers, and character performances. Choose a preset or custom voice, guide delivery with inline audio tags, and adjust voice stability and similarity to shape the final performance.
It is suited to high-quality spoken content where expressive delivery, multilingual speech, and voice consistency matter more than the lower-latency focus of the Turbo variant.
Why Choose This?
-
Expressive text-to-speech
Generate natural-sounding speech for narration, dialogue, and performance-driven content. -
Audio-tag control
Guide delivery with inline cues such as[whispering]and[laughing]. -
Preset and custom voices
Choose a suggested voice or provide a customvoice_id. -
Multilingual speech
Generate spoken content across multiple languages. -
Voice controls
Adjuststabilityandsimilarityto tune delivery consistency and resemblance to the selected voice. -
MP3 output
Receive the generated speech as a downloadable audio file.
Parameters
| Parameter | Required | Description |
|---|---|---|
| text | Yes | Text to convert into speech. Supports 1–10000 characters. Inline audio tags can be included to guide delivery. |
| voice_id | Yes | Suggested voice name or custom voice ID. Default: Alice. |
| stability | No | Controls delivery consistency. Range: 0–1. Default: 0.5. |
| similarity | No | Controls similarity to the selected voice. Range: 0–1. Default: 0.75. |
How to Use
- Enter the text — Provide the script you want spoken.
- Choose a voice — Select a suggested voice or provide a custom voice ID.
- Add expressive cues optional — Use supported audio tags such as
[whispering]or[laughing]where needed. - Adjust voice controls optional — Set
stabilityandsimilarityto tune the performance. - Submit — Generate the speech and retrieve the MP3 audio file.
Pricing
Pricing is $0.08 per 1,000 input characters, prorated according to the actual length of text.
| Input Characters | Price |
|---|---|
| 100 | $0.008 |
| 500 | $0.040 |
| 1,000 | $0.080 |
| 10,000 | $0.800 |
Billable characters include the complete text field, including spaces, punctuation, and inline audio tags.
Requests shorter than 1,000 characters are billed proportionally rather than as a full 1,000-character block.
The calculated charge is rounded up to the nearest $0.000001.
voice_id, stability, and similarity do not add separate charges.
Best Use Cases
- Narration — Generate polished speech for stories, explainers, tutorials, and educational content.
- Voiceovers — Create spoken audio for videos, advertisements, presentations, and social content.
- Character performances — Use expressive tags and voice controls for dialogue and performance-focused speech.
- Multilingual content — Generate spoken content for international and localized workflows.
- Creator content — Produce voice tracks for podcasts, videos, shorts, and other media.
- High-quality speech generation — Use the standard Eleven V4 model when expressive output is the priority.
Pro Tips
- Write scripts in a natural, conversational style for more fluid speech.
- Use expressive audio tags only where the delivery should clearly change.
- Use
stabilityto adjust how consistent the performance remains across the script. - Adjust
similaritywhen you want to tune how closely the result follows the selected voice. - Use punctuation intentionally because it can affect pauses and phrasing.
- Split longer scripts into separate requests when different sections need different voices or delivery styles.
Notes
textsupports1–10000characters.- Spaces, punctuation, and audio tags count toward pricing.
- Requests are billed proportionally by character count.
- Voice availability depends on access to the selected voice.
- Audio tags guide delivery but do not guarantee an exact performance.
- SSML, separate style controls, and speed controls are not exposed by this endpoint.
- The endpoint returns an audio file rather than a live audio stream.
Related Models
- Eleven V4 Turbo — Generate expressive speech with lower-latency text-to-speech generation.
- Eleven V4 — Generate expressive text-to-speech with the standard Eleven V4 model.
Authentication
For authentication details, please refer to the Authentication Guide.
API Endpoints
Submit Task & Query Result
set -euo pipefail
export WAVESPEED_API_KEY="your-api-key"
REQUEST_BODY=$(cat <<'JSON'
{
"text": "[warmly] Welcome to a new day. Take a deep breath, find your rhythm, and make a little room for something wonderful.",
"voice_id": "Alice",
"stability": 0.5,
"similarity": 0.75
}
JSON
)
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/elevenlabs/eleven-v4" \
-H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
-H "Content-Type: application/json" \
-d "${REQUEST_BODY}")
TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body \
"${RESULT_URL}" \
-H "Authorization: Bearer ${WAVESPEED_API_KEY}")
RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')
case "${STATUS}" in
completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
failed|cancelled|timeout|deleted) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
*) sleep 2 ;;
esac
doneParameters
Task Submission Parameters
Request Parameters
| Parameter | Type | Required | Default | Range | Description |
|---|---|---|---|---|---|
| text | string | Yes | - | - | Text to convert to speech. Supports expressive audio tags such as [whispering] and [laughing]. Maximum 10,000 characters for this endpoint. |
| voice_id | string | Yes | Alice | - | Choose a suggested voice or enter a custom voice ID available to the account. |
| stability | number | No | 0.5 | 0 ~ 1 | How consistent the delivery stays. Lower values allow more expressive variation; higher values favor consistency. |
| similarity | number | No | 0.75 | 0 ~ 1 | How closely the generated speech follows the selected voice. Higher values increase voice similarity. |
Response Parameters
| Parameter | Type | Description |
|---|---|---|
| code | integer | HTTP status code (e.g., 200 for success) |
| message | string | Status message (e.g., “success”) |
| data.id | string | Unique identifier for the prediction, Task Id |
| data.model | string | Model ID used for the prediction |
| data.outputs | array | Output values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed) |
| data.urls | object | Object containing related API endpoints |
| data.status | string | Task status. completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses. |
| data.created_at | string | ISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”) |
| data.error | string | Error message (empty if no error occurred) |
| data.timings | object | Object containing timing details |
| data.timings.inference | integer | Inference time in milliseconds |
Result Request Parameters
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
| id | string | Yes | - | Task ID |
Result Response Parameters
| Parameter | Type | Description |
|---|---|---|
| code | integer | HTTP status code (e.g., 200 for success) |
| message | string | Status message (e.g., “success”) |
| data | object | The prediction data object containing all details |
| data.id | string | Unique identifier for the prediction |
| data.model | string | Model ID used for the prediction |
| data.outputs | array<string | object> | Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model. |
| data.urls | object | Object containing related API endpoints |
| data.status | string | Status: completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses |
| data.created_at | string | ISO timestamp of when the request was created |
| data.error | string | Error message (empty if no error occurred) |
| data.timings | object | Object containing timing details |
| data.timings.inference | integer | Inference time in milliseconds |