Seedream 5.0 Flash is LIVE — Faster & Cheaper | Try Now →

google/gemini-3.8-flash/text-to-speech

Gemini 3.8 Flash Text-to-Speech generates expressive speech from text with configurable voice and delivery-style controls for narration, dialogue, localization, virtual assistants, and voice production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

text-to-audio
Input
Enable Safety Checker

Idle

$0.05per run·~20 / $1

ExamplesView all

Related Models

README

Gemini 3.8 Flash Text-to-Speech

Gemini 3.8 Flash Text-to-Speech generates expressive speech from text with 30 preset voices and natural-language control over delivery. Use style instructions to guide tone, pacing, accent, and emotion without embedding delivery directions into the spoken text.

The endpoint supports both standard single-speaker narration and API-based two-speaker dialogue, making it suitable for voiceovers, conversations, character dialogue, narration, and other speech-generation workflows.

Why Choose This?

  • Expressive text-to-speech
    Generate natural speech from text with control over tone, pacing, accent, and emotional delivery.

  • 30 preset voices
    Choose from a broad set of built-in voices for different narration and character styles.

  • Natural-language style control
    Describe how the speech should be delivered through style_instructions.

  • Single-speaker narration
    Generate voiceovers, narration, explanations, and other single-speaker speech.

  • Per-turn delivery control
    Dialogue turns can include their own style_instructions for different emotions or speaking styles.

  • WAV output
    Receive the generated speech as a downloadable WAV audio file.

Parameters

ParameterRequiredDescription
textConditionalText to speak for single-speaker generation. Supports up to 8000 characters. Omit when using dialogue mode.
voiceNoPreset voice for single-speaker speech. Default: Kore.
style_instructionsNoOptional natural-language instructions for tone, pacing, accent, emotion, or other delivery characteristics. Supports up to 2000 characters.
speakersConditionalDefine exactly 2 speakers with distinct speaker_id values and their selected voices.
turnsConditionalEach turn contains a speaker_id, text, and optional style_instructions.

How to Use

Single-Speaker Speech

  1. Enter the text — Provide the content you want spoken.
  2. Choose a voice optional — Select one of the available preset voices or keep the default Kore.
  3. Add style instructions optional — Describe the desired tone, pacing, accent, emotion, or delivery.
  4. Submit — Generate the speech and retrieve the WAV audio.

Two-Speaker Dialogue

  1. Configure two speakers — Define exactly two distinct speaker_id values and assign a voice to each.
  2. Create ordered turns — Add each speaker's dialogue in the order it should be spoken.
  3. Add per-turn style optional — Give individual turns their own delivery instructions when needed.
  4. Add global style optional — Use style_instructions for overall delivery guidance.
  5. Submit — Generate the complete two-speaker dialogue as audio.

Pricing

Billable CharactersPrice
100$0.05
1,000$0.05
5,000$0.25

For single-speaker generation, billable characters include:

  • text
  • style_instructions

For dialogue generation, billable characters include:

  • Text from all turns
  • Per-turn style_instructions
  • Global style_instructions

Best Use Cases

  • Narration and voiceovers — Generate expressive speech for videos, presentations, explainers, and other media.
  • Two-speaker conversations — Create API-driven dialogue between two configured voices.
  • Character dialogue — Give different speakers distinct voices and delivery styles.
  • Audiobook and storytelling workflows — Control pacing, emotion, and narration style through natural-language instructions.
  • Educational content — Generate clear spoken explanations, lessons, and instructional audio.
  • Localized voice content — Guide accent, tone, and delivery through style instructions.
  • Conversational prototypes — Create spoken interactions for assistants, characters, demos, and dialogue experiences.

Pro Tips

  • Keep the spoken content in text or turns and put delivery directions in style_instructions.
  • Use short, specific style instructions for more consistent delivery.
  • Describe concrete characteristics such as pace, emotional tone, energy, or accent instead of vague style requests.
  • In dialogue mode, make sure every speaker_id used in turns has been configured in speakers.
  • Use per-turn style instructions when a speaker's emotion or delivery changes during the conversation.
  • Keep the total request within the 8192-token input limit.
  • Inline vocal events such as <laugh> or <sigh> can be used when they are part of the intended performance.

Related Models

Note:This website uses AI models provided by third parties. Documentation prices are for reference and may be outdated. The Generate button shows an estimate; the final task charge prevails.

Gemini 3.8 Flash Text To Speech API — Quick start

Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/google/gemini-3.8-flash/text-to-speech with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Gemini 3.8 Flash Text To Speech below.

HTTP example
set -euo pipefail

: "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"

REQUEST_BODY=$(cat <<'JSON'
{
    "voice": "Kore"
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/google/gemini-3.8-flash/text-to-speech" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $WAVESPEED_API_KEY" \
  -d "$REQUEST_BODY")

TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
    -H "Authorization: Bearer $WAVESPEED_API_KEY")
  RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
  STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
  case "$STATUS" in
    completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
    failed|cancelled|timeout|deleted) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
    *) sleep 2 ;;
  esac
done
Node.js example
const submitUrl = "https://api.wavespeed.ai/api/v3/google/gemini-3.8-flash/text-to-speech";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');

async function requestJson(url, options = {}) {
  const response = await fetch(url, options);
  if (!response.ok) throw new Error(await response.text());
  return response.json();
}

// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${apiKey}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
        "voice": "Kore"
}),
});
const task = body.data ?? body;
if (!task.id) throw new Error("Submission response did not contain a prediction id");
const resultUrl = `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;

// 2. Poll until the prediction finishes.
while (true) {
  const resultBody = await requestJson(resultUrl, {
    headers: { "Authorization": `Bearer ${apiKey}` },
  });
  const result = resultBody.data ?? resultBody;
  if (result.status === "completed") {
    console.log(result.outputs);
    break;
  }
  if (["failed", "cancelled", "timeout", "deleted"].includes(result.status)) throw new Error(JSON.stringify(result));
  await new Promise(resolve => setTimeout(resolve, 2000));
}
Python example
import json
import os
import time
from urllib.request import Request, urlopen

api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
    "voice": "Kore"
}

def request_json(url, data=None):
    request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
    with urlopen(request) as response:
        return json.load(response)

# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/google/gemini-3.8-flash/text-to-speech", json.dumps(payload).encode())
task = body.get("data", body)
if not task.get("id"):
    raise RuntimeError("Submission response did not contain a prediction id")
result_url = f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"

# 2. Poll until the prediction finishes.
while True:
    result_body = request_json(result_url)
    result = result_body.get("data", result_body)
    status = result.get("status")
    if status == "completed":
        print(result.get("outputs", []))
        break
    if status in {"failed", "cancelled", "timeout", "deleted"}:
        raise RuntimeError(result)
    time.sleep(2)

Gemini 3.8 Flash Text To Speech API — Frequently asked questions

What is the Gemini 3.8 Flash Text To Speech API?

Gemini 3.8 Flash Text To Speech is a Google model for audio generation, exposed as a REST API on WaveSpeedAI. Gemini 3.8 Flash Text-to-Speech generates expressive speech from text with configurable voice and delivery-style controls for narration, dialogue, localization, virtual assistants, and voice production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. You can call it programmatically or try it from the playground above.

How do I call the Gemini 3.8 Flash Text To Speech API?

POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates Python, JavaScript, and cURL examples for submitting requests and polling results. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/google/google-gemini-3.8-flash-text-to-speech.

How much does Gemini 3.8 Flash Text To Speech cost per run?

Gemini 3.8 Flash Text To Speech starts at $0.05 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.

What inputs does Gemini 3.8 Flash Text To Speech accept?

Key inputs: `text`, `voice`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/google/google-gemini-3.8-flash-text-to-speech.

How do I get started with the Gemini 3.8 Flash Text To Speech API?

Sign up for a free WaveSpeedAI account to claim starter credits, copy your API key from /accesskey, then call the endpoint shown in the API tab of the playground. The playground also auto-generates a code sample in Python, JavaScript, or cURL for the parameters you've set.

Can I use Gemini 3.8 Flash Text To Speech outputs commercially?

Commercial usage rights depend on the model's license, set by its provider (Google). Check the provider's applicable terms and WaveSpeedAI's Terms of Service before commercial use.

llms.txt — google/gemini-3.8-flash/text-to-speech for AI agents and LLMs