Nano Banana 2.1 現已上線 — Google 最新模型 | 立即體驗 →

elevenlabs/multilingual-v2

ElevenLabs Multilingual V2 is a multilingual text-to-speech model; cost $0.1 per 1000 characters. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

Text to Speech
輸入
Enable Safety Checker

就緒

$0.2每次運行·~50 / $10

示例查看全部

相關模型

README

ElevenLabs — Multilingual V2 Text-to-Speech

Multilingual V2 converts written text into natural, expressive speech across multiple languages. It delivers clear pronunciation, smooth pacing, and lifelike tone—ideal for voiceovers, narration, learning content, product videos, and global customer support. See the list here.

Key Features

  • High naturalness with humanlike intonation and timing
  • Strong multilingual support and improved accent handling
  • Tunable delivery via similarity and stability
  • Speaker Boost — enhances similarity to the original speaker's voice.

Pricing

Billed by the exact character count of the input text, prorated — no rounding up to 1,000.

  • Rate: $0.20 per 1,000 characters ($200 per 1M characters)
  • No minimum charge; a 100-character request costs $0.020
CharactersCost
100$0.020
500$0.10
1,000$0.20
10,000$2.00

How to Use

  1. Enter your script in the text field.
  2. Choose a voice_id from the built-in catalog or your custom voices. See the voice list for options.
  3. Optional controls • similarity: 0–1 (higher = closer to the base voice timbre) • stability: 0–1 (higher = more consistent delivery) • use_speaker_boost: enhances similarity to the original speaker's voice
  4. Click Run to synthesize and preview your audio.

Notes

  • Use clear punctuation and split very long text into shorter segments for the most stable prosody.

  • voice_id must be valid; if you see a voice-ID error, pick one from the official list linked above.

  • Speaker Boost may slightly increase generation latency.

  • Each request accepts 1–10,000 characters. Split longer text into separate requests.

  • Use <break time="1.5s" /> to add a pause of up to 3 seconds.

提示:本網站部分功能由第三方 AI 模型提供支援。文件價格僅供參考,可能已過時。Generate 按鈕顯示預估價格,最終以任務實際收費為準。

Multilingual v2 API — Quick start

Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/elevenlabs/multilingual-v2 with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Multilingual v2 below.

HTTP example
set -euo pipefail

: "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"

REQUEST_BODY=$(cat <<'JSON'
{
    "text": "A clear example input",
    "voice_id": "Alicia",
    "similarity": 1,
    "stability": 0.5,
    "use_speaker_boost": true
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/elevenlabs/multilingual-v2" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $WAVESPEED_API_KEY" \
  -d "$REQUEST_BODY")

TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
    -H "Authorization: Bearer $WAVESPEED_API_KEY")
  RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
  STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
  case "$STATUS" in
    completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
    failed|cancelled|timeout|deleted) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
    *) sleep 2 ;;
  esac
done
Node.js example
const submitUrl = "https://api.wavespeed.ai/api/v3/elevenlabs/multilingual-v2";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');

async function requestJson(url, options = {}) {
  const response = await fetch(url, options);
  if (!response.ok) throw new Error(await response.text());
  return response.json();
}

// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${apiKey}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
        "text": "A clear example input",
        "voice_id": "Alicia",
        "similarity": 1,
        "stability": 0.5,
        "use_speaker_boost": true
}),
});
const task = body.data ?? body;
if (!task.id) throw new Error("Submission response did not contain a prediction id");
const resultUrl = `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;

// 2. Poll until the prediction finishes.
while (true) {
  const resultBody = await requestJson(resultUrl, {
    headers: { "Authorization": `Bearer ${apiKey}` },
  });
  const result = resultBody.data ?? resultBody;
  if (result.status === "completed") {
    console.log(result.outputs);
    break;
  }
  if (["failed", "cancelled", "timeout", "deleted"].includes(result.status)) throw new Error(JSON.stringify(result));
  await new Promise(resolve => setTimeout(resolve, 2000));
}
Python example
import json
import os
import time
from urllib.request import Request, urlopen

api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
    "text": "A clear example input",
    "voice_id": "Alicia",
    "similarity": 1,
    "stability": 0.5,
    "use_speaker_boost": True
}

def request_json(url, data=None):
    request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
    with urlopen(request) as response:
        return json.load(response)

# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/elevenlabs/multilingual-v2", json.dumps(payload).encode())
task = body.get("data", body)
if not task.get("id"):
    raise RuntimeError("Submission response did not contain a prediction id")
result_url = f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"

# 2. Poll until the prediction finishes.
while True:
    result_body = request_json(result_url)
    result = result_body.get("data", result_body)
    status = result.get("status")
    if status == "completed":
        print(result.get("outputs", []))
        break
    if status in {"failed", "cancelled", "timeout", "deleted"}:
        raise RuntimeError(result)
    time.sleep(2)

Multilingual v2 API — Frequently asked questions

What is the Multilingual v2 API?

Multilingual v2 is a ElevenLabs model for speech synthesis, exposed as a REST API on WaveSpeedAI. ElevenLabs Multilingual V2 is a multilingual text-to-speech model; cost $0.1 per 1000 characters. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. You can call it programmatically or try it from the playground above.

How do I call the Multilingual v2 API?

POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates Python, JavaScript, and cURL examples for submitting requests and polling results. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/elevenlabs/elevenlabs-multilingual-v2.

How much does Multilingual v2 cost per run?

Multilingual v2 starts at $0.2 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.

What inputs does Multilingual v2 accept?

Key inputs: `similarity`, `stability`, `text`, `use_speaker_boost`, `voice_id`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/elevenlabs/elevenlabs-multilingual-v2.

How long does Multilingual v2 take to generate?

Reported generation time on WaveSpeedAI is around 6 seconds per request. This is an estimate, not a latency guarantee; queue time and input settings can change the total wait. live status is visible in the prediction record.

Can I use Multilingual v2 outputs commercially?

Commercial usage rights depend on the model's license, set by its provider (ElevenLabs). Check the provider's applicable terms and WaveSpeedAI's Terms of Service before commercial use.

llms.txt — elevenlabs/multilingual-v2 for AI agents and LLMs