MiniMax H3 오픈 웨이트 | 비디오 생성기에서 사용해보기 →

Inworld Realtime TTS 2 | Realistic Voice & TTS

inworld/

Inworld Realtime TTS-2 converts text into low-latency, natural speech with official TTS-2 controls for delivery mode, language, timestamps, text normalization, and audio output settings. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

text-to-audio
입력

대기 중

$0.035실행당·~28 / $1

예시전체 보기

관련 모델

README

Inworld Realtime TTS 2

Inworld Realtime TTS 2 converts text into natural-sounding speech with low-latency generation and flexible voice controls. It supports multiple output audio formats and lets you adjust speaking rate and temperature for different delivery styles.

Why Choose This?

  • Low-latency text-to-speech Generate speech quickly for interactive apps, assistants, and real-time voice experiences.

  • Natural voice output Create smooth, human-like speech from plain text with selectable voices.

  • Flexible voice controls Adjust speaking rate and temperature to better match tone, pacing, and delivery style.

  • Multiple output formats Export audio in MP3, LINEAR16, OGG_OPUS, FLAC, or WAV depending on your workflow.

  • Production-ready API Access the model through a realtime-friendly API for apps, agents, games, and voice products.

Parameters

ParameterRequiredDescription
textYesInput text to convert into speech.
voice_idNoVoice selection for the generated speech, such as Julia.
speaking_rateNoControls how fast the voice speaks. Default: 1.
temperatureNoControls variation and expressiveness in the generated speech. Default: 1.
output_formatNoOutput audio format: MP3, LINEAR16, OGG_OPUS, FLAC, or WAV.

How to Use

  1. Enter your text — paste or type the content you want to convert into speech.
  2. Choose a voice — select the voice that best fits your use case.
  3. Adjust speaking rate and temperature (optional) — fine-tune pacing and expressiveness.
  4. Choose output format — select MP3, LINEAR16, OGG_OPUS, FLAC, or WAV.
  5. Submit — generate and download the audio output.

Example Input

Welcome to our product demo. Today we will walk through the key features, explain how the workflow operates, and show how quickly you can integrate voice output into your application.

Pricing

Text LengthCost
1–1000 chars$0.035
1001–2000 chars$0.070
2001–3000 chars$0.105
3001–4000 chars$0.140
4001–5000 chars$0.175

Billing Rules

  • Pricing is based on the length of text.
  • Character count is rounded up to the next 1,000-character block.
  • Each additional started 1,000 characters adds $0.035.
  • voice_id, speaking_rate, temperature, and output_format do not affect pricing.

Best Use Cases

  • Realtime voice agents — Generate spoken responses for assistants, NPCs, and conversational interfaces.
  • Interactive applications — Add live voice output to games, education tools, and customer-facing apps.
  • Accessibility features — Turn written content into audio for more accessible user experiences.
  • Content narration — Create voiceovers for guides, product demos, and short-form content.
  • Prototype voice experiences — Quickly test different voices, pacing, and formats in development workflows.

Pro Tips

  • Keep input text clean and well-punctuated for more natural speech rhythm.
  • Split very long content into smaller sections when you want tighter pacing control.
  • Use speaking_rate to match the use case, such as slower for tutorials and faster for assistants.
  • Adjust temperature when you want more variation in delivery style.
  • Choose MP3 for broad compatibility, and use lossless formats like WAV or FLAC when audio quality matters more.
  • Reuse the same voice and settings across related clips for a more consistent user experience.

Notes

  • text is the only required field.
  • Supported output formats are MP3, LINEAR16, OGG_OPUS, FLAC, and WAV.
  • Pricing depends only on text length.
  • Audio format and voice settings do not change the price.

Related Models

  • Other Inworld speech and voice generation models may be useful when you need different latency, quality, or voice configuration options.
참고:이 웹사이트는 제3자가 제공하는 AI 모델을 사용합니다.

Realtime Tts 2 API — Quick start

Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/inworld/realtime-tts-2 with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Realtime Tts 2 below.

HTTP example
set -euo pipefail

: "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"

REQUEST_BODY=$(cat <<'JSON'
{
    "text": "A clear example input",
    "voice_id": "Dennis",
    "speaking_rate": 1,
    "temperature": 1,
    "output_format": "MP3"
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/inworld/realtime-tts-2" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $WAVESPEED_API_KEY" \
  -d "$REQUEST_BODY")

TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL=$(printf '%s' "$TASK" | jq -r '.urls.get // empty')
if [ -z "$RESULT_URL" ]; then
  RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"
fi

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
    -H "Authorization: Bearer $WAVESPEED_API_KEY")
  RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
  STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
  case "$STATUS" in
    completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
    failed|cancelled|timeout) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
    created|processing) sleep 2 ;;
    *) printf 'Unexpected status: %s
' "$STATUS" >&2; exit 1 ;;
  esac
done
Node.js example
const submitUrl = "https://api.wavespeed.ai/api/v3/inworld/realtime-tts-2";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');

async function requestJson(url, options = {}) {
  const response = await fetch(url, options);
  if (!response.ok) throw new Error(await response.text());
  return response.json();
}

// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${apiKey}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
        "text": "A clear example input",
        "voice_id": "Dennis",
        "speaking_rate": 1,
        "temperature": 1,
        "output_format": "MP3"
}),
});
const task = body.data ?? body;
if (!task.id) throw new Error("Submission response did not contain a prediction id");
const resultUrl = task.urls?.get ||
  `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;

// 2. Poll until the prediction finishes.
while (true) {
  const resultBody = await requestJson(resultUrl, {
    headers: { "Authorization": `Bearer ${apiKey}` },
  });
  const result = resultBody.data ?? resultBody;
  if (result.status === "completed") {
    console.log(result.outputs);
    break;
  }
  if (["failed", "cancelled", "timeout"].includes(result.status)) throw new Error(JSON.stringify(result));
  if (!["created", "processing"].includes(result.status)) throw new Error("Unexpected status: " + result.status);
  await new Promise(resolve => setTimeout(resolve, 2000));
}
Python example
import json
import os
import time
from urllib.request import Request, urlopen

api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
    "text": "A clear example input",
    "voice_id": "Dennis",
    "speaking_rate": 1,
    "temperature": 1,
    "output_format": "MP3"
}

def request_json(url, data=None):
    request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
    with urlopen(request) as response:
        return json.load(response)

# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/inworld/realtime-tts-2", json.dumps(payload).encode())
task = body.get("data", body)
if not task.get("id"):
    raise RuntimeError("Submission response did not contain a prediction id")
result_url = task.get("urls", {}).get("get") or f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"

# 2. Poll until the prediction finishes.
while True:
    result_body = request_json(result_url)
    result = result_body.get("data", result_body)
    status = result.get("status")
    if status == "completed":
        print(result.get("outputs", []))
        break
    if status in {"failed", "cancelled", "timeout"}:
        raise RuntimeError(result)
    if status not in {"created", "processing"}:
        raise RuntimeError(f"Unexpected status: {status}")
    time.sleep(2)

Realtime Tts 2 API — Frequently asked questions

What is the Realtime Tts 2 API?

Realtime Tts 2 is a Inworld model for audio generation, exposed as a REST API on WaveSpeedAI. Inworld Realtime TTS-2 converts text into low-latency, natural speech with official TTS-2 controls for delivery mode, language, timestamps, text normalization, and audio output settings. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. You can call it programmatically or try it from the playground above.

How do I call the Realtime Tts 2 API?

POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates production-oriented Python, JavaScript, and cURL examples with timeouts, transient-error handling, and safe GET retries. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/inworld/inworld-realtime-tts-2.

How much does Realtime Tts 2 cost per run?

Realtime Tts 2 starts at $0.035 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.

What inputs does Realtime Tts 2 accept?

Key inputs: `enable_sync_mode`, `output_format`, `speaking_rate`, `temperature`, `text`, `voice_id`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/inworld/inworld-realtime-tts-2.

How long does Realtime Tts 2 take to generate?

Median end-to-end generation time on WaveSpeedAI is around 2 seconds per request, based on recent successful runs. Queue time varies with global demand; live status is visible in the prediction record.

Can I use Realtime Tts 2 outputs commercially?

Commercial usage rights depend on the model's license, set by its provider (Inworld). The license summary appears on the model card above; see WaveSpeedAI's Terms of Service for platform-level conditions.

Inworld Realtime TTS 2 | Realistic Voice & TTS API on WaveSpeedAI