Nano Banana 2.1 ya DISPONIBLE — Lo último de Google | Pruébalo →
Speech Generation

Speech Generation

Convert text into expressive spoken audio

Nuestra selección

chatterbox/speech-to-speech
speech-to-speech

chatterbox/speech-to-speech

Chatterbox Speech to Speech is a fast AI voice conversion model that converts source audio into a target voice style with optional reference audio guidance. Ready-to-use REST inference API for voice conversion, speech style transfer, dubbing, character voices, creator content, audio localization, and professional speech-to-speech workflows with simple integration, no coldstarts, and affordable pricing.

Todos los modelos

40 modelos
chatterbox/speech-to-speech
speech-to-speech

chatterbox/speech-to-speech

Chatterbox Speech to Speech is a fast AI voice conversion model that converts source audio into a target voice style with optional reference audio guidance. Ready-to-use REST inference API for voice conversion, speech style transfer, dubbing, character voices, creator content, audio localization, and professional speech-to-speech workflows with simple integration, no coldstarts, and affordable pricing.

chatterbox/text-to-speech
text-to-speech

chatterbox/text-to-speech

Chatterbox Text to Speech is a fast AI TTS model that converts text into expressive speech with optional reference audio, emotive tags, and delivery controls. Ready-to-use REST inference API for voice generation, narration, character dialogue, dubbing, virtual assistants, creator content, and professional text-to-speech workflows with simple integration, no coldstarts, and affordable pricing.

bytedance/seed-speech-tts-2.0
text-to-speech

bytedance/seed-speech-tts-2.0

ByteDance Seed Speech TTS 2.0 is a fast AI text-to-speech model that converts text into natural speech with multilingual voices, delivery controls, and MP3 or Opus output. Ready-to-use REST inference API for voice generation, narration, dubbing, virtual assistants, product demos, creator content, and professional TTS workflows with simple integration, no coldstarts, and affordable pricing.

kwaivgi/kling-text-to-audio
sound-effects

kwaivgi/kling-text-to-audio

Kling Text-to-Audio turns text prompts into custom sound effects for videos, games, and multimedia using KlingAI's audio model. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

inworld/realtime-tts-2
text-to-speech

inworld/realtime-tts-2

Inworld Realtime TTS-2 converts text into low-latency, natural speech with official TTS-2 controls for delivery mode, language, timestamps, text normalization, and audio output settings. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-3.1-flash/text-to-speech
text-to-speech

google/gemini-3.1-flash/text-to-speech

Gemini 3.1 Flash Text to Speech generates expressive multi-speaker audio from text, with natural voices and multilingual language control for dialogue, narration, localization, and AI voice workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

minimax/speech-2.6-turbo
text-to-speech

minimax/speech-2.6-turbo

Minimax Speech 2.6 Turbo is a Text-to-Speech model offering ultra-human voice cloning, industry-leading text normalization, sub-250ms latency and 40+ language support. Pricing: $0.06 per 1000 characters. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

minimax/speech-2.8-turbo
text-to-speech

minimax/speech-2.8-turbo

MiniMax Speech 2.8 Turbo is a high-definition text-to-speech model with natural and expressive voice synthesis. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

elevenlabs/turbo-v2
text-to-speech

elevenlabs/turbo-v2

ElevenLabs Turbo V2 is a Text-To-Speech model available via WaveSpeedAI, billed at $0.05 per 1000 characters for API requests. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

elevenlabs/flash-v2
text-to-speech

elevenlabs/flash-v2

ElevenLabs Flash V2 is a Text-to-Speech model that converts text into spoken audio using the ElevenLabs Flash V2 engine. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

elevenlabs/flash-v2.5
text-to-speech

elevenlabs/flash-v2.5

ElevenLabs Flash v2.5 is a text-to-speech model on WaveSpeedAI, billed at $0.05 per 1000 characters for generated speech. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

kwaivgi/kling-v1-tts
text-to-speech

kwaivgi/kling-v1-tts

Kling V1 TTS creates natural-sounding audio and supports KlingAI image, video, sound effect, virtual model, and custom AI workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

alibaba/qwen3-tts-flash
text-to-speech

alibaba/qwen3-tts-flash

Qwen3 TTS Flash: Low-latency Text-to-Speech for English and Chinese with multiple voices, ideal for real-time dialogue. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

elevenlabs/eleven-v3
text-to-speech

elevenlabs/eleven-v3

ElevenLabs eleven-v3 is a text-to-speech model available as a hosted endpoint. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

wavespeed-ai/ace-step
ai-music

wavespeed-ai/ace-step

ACE-Step generates up to 4-minute music with lyrics from text and high acoustic fidelity; supports voice cloning, lyric edits, and remixing. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

wavespeed-ai/ace-step/audio-to-audio
music-editing

wavespeed-ai/ace-step/audio-to-audio

ACE-Step Audio-to-Audio turns existing tracks into remixes or vocal edits using remix and lyrics modes while preserving audio character. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

wavespeed-ai/ace-step/prompt-to-audio
ai-music

wavespeed-ai/ace-step/prompt-to-audio

ACE-Step Prompt-to-Audio creates music from simple prompts, auto-generating genre tags and lyrics for quick song creation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

wavespeed-ai/ace-step/audio-outpaint
music-editing

wavespeed-ai/ace-step/audio-outpaint

ACE-Step Audio Outpaint generates seamless start or end extensions that match the original, ideal for intros, outros and longer tracks. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

wavespeed-ai/ace-step/audio-inpaint
music-editing

wavespeed-ai/ace-step/audio-inpaint

ACE-Step Audio Inpaint edits a specific audio segment to change lyrics or style while preserving the surrounding audio. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

minimax/speech-2.6-hd
text-to-speech

minimax/speech-2.6-hd

Minimax Speech 2.6 HD: Ultra-human, low-latency (< 250ms) TTS with voice cloning, text normalization and support for 40+ languages. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

minimax/speech-02-hd
text-to-speech

minimax/speech-02-hd

Minimax Speech 02 HD is Minimax's high-definition text-to-speech model delivering clear HD voices; pricing $0.05 per 1,000 characters. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

elevenlabs/turbo-v2.5
text-to-speech

elevenlabs/turbo-v2.5

ElevenLabs Turbo V2.5 is a text-to-speech model available via WaveSpeedAI, billed at $0.05 per 1000 characters for TTS requests. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

wavespeed-ai/vibevoice
text-to-speech

wavespeed-ai/vibevoice

wavespeed-ai/vibevoice is an advanced voice generation model for producing high-fidelity, natural, and expressive speech from text, with optional speaker/region-style control for more precise results and easy integration into real-world applications. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

elevenlabs/multilingual-v2
text-to-speech

elevenlabs/multilingual-v2

ElevenLabs Multilingual V2 is a multilingual text-to-speech model; cost $0.1 per 1000 characters. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

minimax/speech-2.8-hd
text-to-speech

minimax/speech-2.8-hd

MiniMax Speech 2.8 HD is a high-definition text-to-speech model with natural and expressive voice synthesis for premium audio quality. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

wavespeed-ai/qwen3-tts/text-to-speech
text-to-speech

wavespeed-ai/qwen3-tts/text-to-speech

Qwen3 TTS: Multi-language, multi-voice text-to-speech synthesis with style control. Supports 11 languages and 9 voice characters. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

wavespeed-ai/qwen3-tts/voice-clone
text-to-speech

wavespeed-ai/qwen3-tts/voice-clone

Qwen3 TTS Voice Clone: Clone any voice from a reference audio and generate speech in that voice. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

wavespeed-ai/qwen3-tts/voice-design
text-to-speech

wavespeed-ai/qwen3-tts/voice-design

Qwen3 TTS Voice Design: Generate speech with custom voice characteristics described in natural language. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.

microsoft/vibevoice
text-to-speech

microsoft/vibevoice

Microsoft VibeVoice text-to-speech model generates long-form speech from text with multi-speaker dialogue support. Choose from 9 voice presets across English, Chinese, and Hindi. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

inworld/inworld-1.5-max/text-to-speech
text-to-speech

inworld/inworld-1.5-max/text-to-speech

Inworld 1.5 Max delivers premium text-to-speech synthesis with 56+ multilingual voices, adjustable speaking rate, and high-fidelity natural-sounding audio output. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

inworld/inworld-1.5-mini/text-to-speech
text-to-speech

inworld/inworld-1.5-mini/text-to-speech

Inworld 1.5 Mini delivers high-quality text-to-speech synthesis with 56+ multilingual voices, adjustable speaking rate, and natural-sounding audio output. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-2.5-pro/text-to-speech
text-to-speech

google/gemini-2.5-pro/text-to-speech

Google Gemini 2.5 Pro Text-to-Speech delivers natural multi-speaker voice synthesis with 30+ voices across 24 languages. Perfect for dialogues, conversations, and multilingual content. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

google/gemini-2.5-flash/text-to-speech
text-to-speech

google/gemini-2.5-flash/text-to-speech

Google Gemini 2.5 Flash Text-to-Speech delivers fast, natural multi-speaker voice synthesis with 30+ voices across 24 languages at lower cost. Perfect for dialogues, conversations, and multilingual content. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

wavespeed-ai/omnivoice/text-to-speech
text-to-speech

wavespeed-ai/omnivoice/text-to-speech

OmniVoice is a massively multilingual zero-shot TTS supporting 600+ languages. Generate speech with auto voice or design custom voices using natural language descriptions. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

wavespeed-ai/omnivoice/voice-clone
text-to-speech

wavespeed-ai/omnivoice/voice-clone

OmniVoice Voice Clone clones any voice from a short 3-10 second audio sample. Supports 600+ languages with zero-shot voice cloning. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

minimax/speech-2.5-turbo-preview
text-to-speech

minimax/speech-2.5-turbo-preview

Minimax Speech 2.5 Turbo Preview: HD TTS with multilingual support, accurate voice replication across 40 languages. $0.04/1000 chars. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

minimax/speech-2.5-hd-preview
text-to-speech

minimax/speech-2.5-hd-preview

MiniMax Speech 2.5 HD Preview offers HD TTS with enhanced multilingual expressiveness, accurate voice cloning, and 40-language support. Ready-to-use REST API, best performance, no coldstarts, affordable pricing.

minimax/voice-design
text-to-speech

minimax/voice-design

MiniMax Voice Design generates natural voices from textual descriptions - no cloning - lets you set tone, accent and personality. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

minimax/voice-clone
voice-cloning

minimax/voice-clone

Minimax Voice Clone creates high-quality voice clones from short reference clips, closely matching tone, accent, and speaking style. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

minimax/speech-02-turbo
text-to-speech

minimax/speech-02-turbo

Minimax Speech-02 Turbo is a high-definition text-to-speech model delivering natural voice output. Cost: $0.03 per 1000 characters. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

API de Speech Generation — precios y rendimiento

Ejecuta cualquier modelo de la colección Speech Generation a través de una sola API REST. Paga por generación — sin suscripciones ni mínimos — con latencia líder del sector sobre una infraestructura con 99,9 % de disponibilidad.

Por qué ejecutar Speech Generation en WaveSpeedAI

Precios transparentes

Precio por llamada para cada modelo Speech Generation. El precio aparece en la página de cada modelo — sin recargos de plataforma.

Optimizado para baja latencia

La mayoría de los modelos de imagen Speech Generation terminan en menos de 2 segundos. Los modelos de vídeo y 3D son varias veces más rápidos que las alternativas autoalojadas.

99,9 % de disponibilidad

Conmutación por error multirregión y reintentos automáticos mantienen tu tráfico de producción en línea — incluso durante caídas del proveedor.

Preguntas frecuentes

¿Cuánto cuesta la API de Speech Generation?+

Cada modelo tiene su propio precio por llamada listado en su página. Cobramos por generación exitosa, sin cuotas de suscripción ni mínimos.

¿Qué tan rápidos son los modelos Speech Generation en WaveSpeedAI?+

Los modelos de imagen de esta colección suelen completarse en menos de 2 segundos. Los modelos de vídeo y 3D dependen de la duración y la resolución, pero suelen ser varias veces más rápidos que las ejecuciones autoalojadas.

¿Puedo probar la API sin tarjeta de crédito?+

Las cuentas nuevas que cumplan los requisitos pueden recibir $1 en créditos promocionales para probar modelos Speech Generation sin tarjeta de crédito. No se garantizan créditos de prueba en cada registro; consulta tu saldo antes de generar.

¿Hay límites de tasa?+

Las cuentas estándar tienen límites generosos de trabajos concurrentes. Los planes Enterprise ofrecen RPM personalizado, mayor concurrencia y capacidad dedicada — contacta con ventas para más detalles.