Seedream 5.0 Flash is LIVE — Faster & Cheaper | Try Now →

google/gemini-3.8-flash-lite/text-to-speech

Gemini 3.8 Flash-Lite Text-to-Speech generates expressive speech from text with configurable voice and delivery-style controls for narration, dialogue, localization, virtual assistants, and scalable voice production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

text-to-audio
Input
Enable Safety Checker

Idle

$0.04per run·~25 / $1

ExamplesView all

Related Models

README

Gemini 3.8 Flash-Lite Text-to-Speech

Gemini 3.8 Flash-Lite Text-to-Speech generates expressive WAV speech from text with 30 preset voices and natural-language control over tone, pacing, accent, and emotion.

It supports both single-speaker narration and two-speaker dialogue, making it suitable for voiceovers, conversations, storytelling, and other speech-generation workflows.

Why Choose This?

  • Expressive text-to-speech
    Generate natural speech with control over tone, pace, accent, and emotion.

  • 30 preset voices
    Choose from a wide range of built-in voices for different speaking styles.

  • Natural-language delivery control
    Use style_instructions to describe how the speech should sound.

  • Single-speaker narration
    Create voiceovers, narration, explanations, and spoken content.

  • WAV output
    Receive generated speech as a downloadable WAV audio file.

Parameters

ParameterRequiredDescription
textConditionalText to speak for single-speaker generation. Supports up to 8000 characters. Omit when using dialogue mode.
voiceNoPreset voice for single-speaker speech. Default: Kore.
style_instructionsNoOptional instructions for tone, pacing, accent, emotion, and delivery style. Supports up to 2000 characters.
speakersConditionalConfiguration for two dialogue speakers and their voices.
turnsConditionalOrdered dialogue turns containing the speaker, text, and optional delivery instructions.

How to Use

  1. Enter the text — Provide the content you want spoken.
  2. Choose a voice — Select one of the available preset voices.
  3. Add style instructions optional — Describe the desired tone, pace, accent, or emotion.
  4. Submit — Generate and retrieve the WAV audio.

For dialogue, configure two speakers and provide their lines through turns.

Pricing

Pricing is $0.04 per 1,000 billable characters, with a minimum charge of $0.04 per request.

Billable CharactersPrice
100$0.04
1,000$0.04
5,000$0.20

Text and delivery instructions contribute to billable character count.

Best Use Cases

  • Voiceovers — Generate speech for videos, presentations, and promotional content.
  • Narration — Create spoken content for stories, explainers, and educational material.
  • Two-speaker conversations — Generate dialogue between two different preset voices.
  • Character voices — Use voice and style controls for different speaking personalities.
  • Localized content — Guide accent, tone, and delivery for different language workflows.
  • Conversational prototypes — Create spoken interactions for demos, assistants, and characters.

Pro Tips

  • Keep spoken content separate from delivery instructions.
  • Use short, specific style instructions for more consistent results.
  • Describe tone, pace, energy, or emotion directly.
  • Use different voices for clearer separation in dialogue.
  • Inline vocal events such as <laugh> or <sigh> can be used when needed.

Notes

  • text supports up to 8000 characters.
  • style_instructions supports up to 2000 characters.
  • The default voice is Kore.
  • Dialogue mode supports two speakers.

Related Models

Note:This website uses AI models provided by third parties. Documentation prices are for reference and may be outdated. The Generate button shows an estimate; the final task charge prevails.

Gemini 3.8 Flash Lite Text To Speech API — Quick start

Grab a WaveSpeedAI API key, then call POST https://api.wavespeed.ai/api/v3/google/gemini-3.8-flash-lite/text-to-speech with your input as JSON. The endpoint returns a prediction id. Start polling the result endpoint around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. On completed, read output values from data.outputs. Examples for Gemini 3.8 Flash Lite Text To Speech below.

HTTP example
set -euo pipefail

: "${WAVESPEED_API_KEY:?Set WAVESPEED_API_KEY}"

REQUEST_BODY=$(cat <<'JSON'
{
    "voice": "Kore"
}
JSON
)

# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  -X POST "https://api.wavespeed.ai/api/v3/google/gemini-3.8-flash-lite/text-to-speech" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $WAVESPEED_API_KEY" \
  -d "$REQUEST_BODY")

TASK=$(printf '%s' "$SUBMIT_RESPONSE" | jq 'if has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "$TASK" | jq -r '.id')
if [ -z "$PREDICTION_ID" ] || [ "$PREDICTION_ID" = "null" ]; then
  printf 'Submission response did not contain a prediction id
' >&2
  exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/$PREDICTION_ID/result"

# 2. Poll until the prediction finishes.
while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body "$RESULT_URL" \
    -H "Authorization: Bearer $WAVESPEED_API_KEY")
  RESULT=$(printf '%s' "$RESPONSE" | jq 'if has("data") then .data else . end')
  STATUS=$(printf '%s' "$RESULT" | jq -r '.status')
  case "$STATUS" in
    completed) printf '%s\n' "$RESULT" | jq '.outputs'; break ;;
    failed|cancelled|timeout|deleted) printf '%s\n' "$RESULT" | jq . >&2; exit 1 ;;
    *) sleep 2 ;;
  esac
done
Node.js example
const submitUrl = "https://api.wavespeed.ai/api/v3/google/gemini-3.8-flash-lite/text-to-speech";
const apiKey = process.env.WAVESPEED_API_KEY;
if (!apiKey) throw new Error('Set WAVESPEED_API_KEY');

async function requestJson(url, options = {}) {
  const response = await fetch(url, options);
  if (!response.ok) throw new Error(await response.text());
  return response.json();
}

// 1. Submit the prediction.
const body = await requestJson(submitUrl, {
  method: "POST",
  headers: {
    "Authorization": `Bearer ${apiKey}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
        "voice": "Kore"
}),
});
const task = body.data ?? body;
if (!task.id) throw new Error("Submission response did not contain a prediction id");
const resultUrl = `https://api.wavespeed.ai/api/v3/predictions/${task.id}/result`;

// 2. Poll until the prediction finishes.
while (true) {
  const resultBody = await requestJson(resultUrl, {
    headers: { "Authorization": `Bearer ${apiKey}` },
  });
  const result = resultBody.data ?? resultBody;
  if (result.status === "completed") {
    console.log(result.outputs);
    break;
  }
  if (["failed", "cancelled", "timeout", "deleted"].includes(result.status)) throw new Error(JSON.stringify(result));
  await new Promise(resolve => setTimeout(resolve, 2000));
}
Python example
import json
import os
import time
from urllib.request import Request, urlopen

api_key = os.environ["WAVESPEED_API_KEY"]
headers = {"Authorization": f"Bearer {api_key}", "Content-Type": "application/json"}
payload = {
    "voice": "Kore"
}

def request_json(url, data=None):
    request = Request(url, data=data, headers=headers, method="POST" if data else "GET")
    with urlopen(request) as response:
        return json.load(response)

# 1. Submit the prediction.
body = request_json("https://api.wavespeed.ai/api/v3/google/gemini-3.8-flash-lite/text-to-speech", json.dumps(payload).encode())
task = body.get("data", body)
if not task.get("id"):
    raise RuntimeError("Submission response did not contain a prediction id")
result_url = f"https://api.wavespeed.ai/api/v3/predictions/{task['id']}/result"

# 2. Poll until the prediction finishes.
while True:
    result_body = request_json(result_url)
    result = result_body.get("data", result_body)
    status = result.get("status")
    if status == "completed":
        print(result.get("outputs", []))
        break
    if status in {"failed", "cancelled", "timeout", "deleted"}:
        raise RuntimeError(result)
    time.sleep(2)

Gemini 3.8 Flash Lite Text To Speech API — Frequently asked questions

What is the Gemini 3.8 Flash Lite Text To Speech API?

Gemini 3.8 Flash Lite Text To Speech is a Google model for audio generation, exposed as a REST API on WaveSpeedAI. Gemini 3.8 Flash-Lite Text-to-Speech generates expressive speech from text with configurable voice and delivery-style controls for narration, dialogue, localization, virtual assistants, and scalable voice production workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing. You can call it programmatically or try it from the playground above.

How do I call the Gemini 3.8 Flash Lite Text To Speech API?

POST your input parameters to the model's REST endpoint (shown in the API tab of this playground) with your WaveSpeedAI API key in the Authorization header. Submission returns a prediction ID. Poll the result endpoint starting around every 2 seconds, increase the interval for long-running tasks, and stop on any terminal status. The playground generates Python, JavaScript, and cURL examples for submitting requests and polling results. Full request/response shape is documented at https://wavespeed.ai/docs/docs-api/google/google-gemini-3.8-flash-lite-text-to-speech.

How much does Gemini 3.8 Flash Lite Text To Speech cost per run?

Gemini 3.8 Flash Lite Text To Speech starts at $0.04 per run. That figure is the base price — the final charge scales with the parameters you set in the form (output size, length, count, references, or whatever knobs this model exposes), so a higher-quality or larger output costs more than a minimal one. The exact cost for your current input is shown live next to the Generate button before you submit, and the actual per-call charge is recorded on the prediction afterwards.

What inputs does Gemini 3.8 Flash Lite Text To Speech accept?

Key inputs: `text`, `voice`. The full JSON schema (types, defaults, allowed values) is rendered above the Generate button and mirrored in the API reference at https://wavespeed.ai/docs/docs-api/google/google-gemini-3.8-flash-lite-text-to-speech.

How do I get started with the Gemini 3.8 Flash Lite Text To Speech API?

Sign up for a free WaveSpeedAI account to claim starter credits, copy your API key from /accesskey, then call the endpoint shown in the API tab of the playground. The playground also auto-generates a code sample in Python, JavaScript, or cURL for the parameters you've set.

Can I use Gemini 3.8 Flash Lite Text To Speech outputs commercially?

Commercial usage rights depend on the model's license, set by its provider (Google). Check the provider's applicable terms and WaveSpeedAI's Terms of Service before commercial use.

llms.txt — google/gemini-3.8-flash-lite/text-to-speech for AI agents and LLMs