ByteDance Seed Speech TTS 2.0 is a fast AI text-to-speech model that converts text into natural speech with multilingual voices, delivery controls, and MP3 or Opus output. Ready-to-use REST inference API for voice generation, narration, dubbing, virtual assistants, product demos, creator content, and professional TTS workflows with simple integration, no coldstarts, and affordable pricing.
대기 중
$0.03실행당·~33 / $1
ByteDance Seed Speech TTS 2.0 converts text into speech with a wide selection of multilingual voice presets and controls for language, speed, pitch, volume, sample rate, and output format. It is suitable for narration, voiceovers, character voices, multilingual content, and production-ready speech synthesis workflows.
High-quality text-to-speech Generate natural-sounding speech from plain text.
Large voice preset library Choose from many built-in voices across English, Chinese, Japanese, Spanish, Indonesian, Portuguese, Korean, Italian, German, and French.
Multilingual support Use a language override or leave it empty for automatic language detection.
Fine-grained voice controls Adjust speed, volume, pitch, sample rate, and optional voice instructions.
Voice instruction support Add natural-language instructions for tone, emotion, pace, or volume without having that instruction spoken aloud.
Production-ready API Suitable for narration, audiobooks, short-form content, virtual assistants, localization, and voice-based creative workflows.
| Parameter | Required | Description |
|---|---|---|
| text | Yes | The text to synthesize into speech. |
| voice | No | Voice preset to use for speech synthesis. Default: stokie_en. |
| voice_instruction | No | Optional natural-language instruction for tone, emotion, pace, or volume. It is not spoken aloud. |
| output_format | No | Output audio format. Supported values: mp3, opus. Default: mp3. |
| sample_rate | No | Sample rate of the output audio in Hz. Supported values: 8000, 16000, 22050, 24000, 32000, 44100, 48000. Default: 24000. |
| speed | No | Speech speed. Range: 0.5–2. Default: 1. |
| volume | No | Speech volume. Range: 0.5–2. Default: 1. |
| pitch | No | Voice pitch shift in semitones. Range: -12–12. Default: 0. |
| language | No | Optional language override. Leave unset for automatic language detection. |
mp3 or opus.Warm, calm, confident narration with a slightly slower pace and soft expressive tone.
Pricing is based on the length of the input text.
| Text Length | Cost |
|---|---|
| 1–1000 chars | $0.03 |
| 1001–2000 chars | $0.06 |
| 2001–3000 chars | $0.09 |
| 3001–4000 chars | $0.12 |
| 4001–5000 chars | $0.15 |
voice, voice_instruction, output_format, sample_rate, speed, volume, pitch, and language do not affect pricingvoice_instruction when you want more expressive control without changing the text itself.speed near 1 for natural speech, then adjust only if needed.language when auto-detection may be ambiguous.text is required.voice_instruction affects delivery style but is not spoken aloud.language can be left empty for automatic language detection.Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/bytedance/seed-speech-tts-2.0 with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Seed Speech Tts 2.0 below.
set -euo pipefail
: "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"
REQUEST_BODY=$(cat <<'JSON'
{
"text": "A clear example input",
"voice": "stokie_en",
"output_format": "mp3",
"sample_rate": 24000,
"speed": 1,
"volume": 1,
"pitch": 0,
"language": "zh"
}
JSON
)
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/bytedance/seed-speech-tts-2.0" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $WAVESPEED_API_KEY" \
-d "$REQUEST_BODY")
TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL=$(printf '%s' "$TASK" | jq -r '.urls.get // empty')
if [ -z "$RESULT_URL" ]; then
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
fi
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
-H "Authorization: Bearer $WAVESPEED_API_KEY")
RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
case "$STATUS" in
completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
failed|cancelled|timeout) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
created|processing) sleep 2 ;;
*) printf 'Unexpected status: %s
' "$STATUS" >&2; exit 1 ;;
esac
doneconst submitUrl = "https://api.wavespeed.ai/api/v3/bytedance/seed-speech-tts-2.0";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');
async function requestJson(url, options = {}) {
const response = await fetch(url, options);
if (!response.ok) throw new Error(await response.text());
return response.json();
}
// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
method: "POST",
headers: {
"Authorization": `Bearer ${apiKey}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
"text": "A clear example input",
"voice": "stokie_en",
"output_format": "mp3",
"sample_rate": 24000,
"speed": 1,
"volume": 1,
"pitch": 0,
"language": "zh"
}),
});
const task = body.data ?? body;
if (!task.id) throw new Error("Submission response did not contain a prediction id");
const resultUrl = task.urls?.get ||
`https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;
// 2. Poll until the prediction finishes.
while (true) {
const resultBody = await requestJson(resultUrl, {
headers: { "Authorization": `Bearer ${apiKey}` },
});
const result = resultBody.data ?? resultBody;
if (result.status === "completed") {
console.log(result.outputs);
break;
}
if (["failed", "cancelled", "timeout"].includes(result.status)) throw new Error(JSON.stringify(result));
if (!["created", "processing"].includes(result.status)) throw new Error("Unexpected status: " + result.status);
await new Promise(resolve => setTimeout(resolve, 2000));
}import json
import os
import time
from urllib.request import Request, urlopen
api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
"text": "A clear example input",
"voice": "stokie_en",
"output_format": "mp3",
"sample_rate": 24000,
"speed": 1,
"volume": 1,
"pitch": 0,
"language": "zh"
}
def request_json(url, data=None):
request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
with urlopen(request) as response:
return json.load(response)
# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/bytedance/seed-speech-tts-2.0", json.dumps(payload).encode())
task = body.get("data", body)
if not task.get("id"):
raise RuntimeError("Submission response did not contain a prediction id")
result_url = task.get("urls", {}).get("get") or f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"
# 2. Poll until the prediction finishes.
while True:
result_body = request_json(result_url)
result = result_body.get("data", result_body)
status = result.get("status")
if status == "completed":
print(result.get("outputs", []))
break
if status in {"failed", "cancelled", "timeout"}:
raise RuntimeError(result)
if status not in {"created", "processing"}:
raise RuntimeError(f"Unexpected status: {status}")
time.sleep(2)Seed Speech Tts 2.0 is a ByteDance model for audio generation, exposed as a REST API on WaveSpeedAI. ByteDance Seed Speech TTS 2.0 is a fast AI text-to-speech model that converts text into natural speech with multilingual voices, delivery controls, and MP3 or Opus output. Ready-to-use REST inference API for voice generation, narration, dubbing, virtual assistants, product demos, creator content, and professional TTS workflows with simple integration, no coldstarts, and affordable pricing. You can call it programmatically or try it from the playground above.
POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates production-oriented Python, JavaScript, and cURL examples with timeouts, transient-error handling, and safe GET retries. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/bytedance/bytedance-seed-speech-tts-2.0.
Seed Speech Tts 2.0 starts at $0.030 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.
Key inputs: `language`, `output_format`, `pitch`, `sample_rate`, `speed`, `text`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/bytedance/bytedance-seed-speech-tts-2.0.
Median end-to-end generation time on WaveSpeedAI is around 7 seconds per request, based on recent successful runs. Queue time varies with global demand; live status is visible in the prediction record.
Commercial usage rights depend on the model's license, set by its provider (ByteDance). The license summary appears on the model card above; see WaveSpeedAI's Terms of Service for platform-level conditions.