Qwen3-TTS Text-to-Speech
Qwen3-TTS Text-to-Speech is a high-quality text-to-speech model with a curated selection of preset voices. Choose from 9 distinct voices spanning different genders and speaking styles, with optional style instructions to fine-tune the delivery.
Why Choose This?
-
Curated voice library
9 preset voices with distinct personalities — from professional narrators to friendly conversational tones.
-
Style instruction support
Guide the speaking style with natural language instructions for customized delivery.
-
Auto language detection
Set language to "auto" and the model intelligently detects the language from your text.
-
Simple and fast
Straightforward interface — select a voice, enter text, and generate.
Parameters
| Parameter | Required | Description |
|---|
| text | Yes | The text to convert to speech |
| language | Yes | Language code or "auto" for automatic detection |
| voice | Yes | Preset voice to use (see Available Voices below) |
| style_instruction | No | Natural language guidance for speaking style |
Available Voices
| Voice | Description |
|---|
| Vivian | Female voice |
| Serena | Female voice |
| Ono_Anna | Female voice |
| Sohee | Female voice |
| Uncle_Fu | Male voice |
| Dylan | Male voice |
| Eric | Male voice |
| Ryan | Male voice |
| Aiden | Male voice |
Style Instruction Examples
- "Speak slowly and calmly, like a meditation guide"
- "Energetic and enthusiastic, like a sports announcer"
- "Professional and clear, suitable for corporate presentations"
- "Warm and friendly, like talking to a close friend"
How to Use
- Enter your text — write or paste the content you want to convert to speech.
- Select language — choose the target language or use "auto" for automatic detection.
- Choose a voice — select from the 9 available preset voices.
- Add style instruction (optional) — describe how you want the voice to sound.
- Run — submit and download your audio file.
Pricing
| Text Length | Cost |
|---|
| Under 1,000 chars | $0.02 |
| 1,000+ chars | $0.02 per 1,000 characters |
Billing Rules
- Minimum charge: $0.02 (for texts under 1,000 characters)
- For longer texts: $0.02 × (character count / 1,000)
Best Use Cases
- Video Voiceovers — Generate professional narration for YouTube, ads, or explainer videos.
- Audiobook Production — Convert manuscripts into natural-sounding narration.
- Podcasts & Broadcasting — Create consistent voice content without recording equipment.
- E-learning & Training — Produce clear, engaging audio for educational materials.
- Accessibility — Convert written content to audio for visually impaired users.
Pro Tips
- Try different voices to find the best match for your content type.
- Use style_instruction to adjust tone without changing the voice itself.
- Match female voices (Vivian, Serena, Ono_Anna, Sohee) for softer content; male voices (Uncle_Fu, Dylan, Eric, Ryan, Aiden) for authoritative content.
- Test with short text first to preview how the voice sounds before generating longer content.
Related Models
- Qwen3-TTS Voice Clone — Clone any voice from a short audio sample.
- Qwen3-TTS Voice Design — Design custom voices using natural language descriptions.
Notes
- All 9 voices are optimized for natural, clear speech output.
- Style instructions work best when they describe emotion, pace, or tone rather than technical audio settings.
- For best quality, match the language parameter to your text content.