/filters:quality(82)/media/images/1773962750571484976_JuAKU3bj.webp)
Convert text into expressive spoken audio
/filters:quality(82)/media/images/1773962750571484976_JuAKU3bj.webp)
Chatterbox Speech to Speech is a fast AI voice conversion model that converts source audio into a target voice style with optional reference audio guidance. Ready-to-use REST inference API for voice conversion, speech style transfer, dubbing, character voices, creator content, audio localization, and professional speech-to-speech workflows with simple integration, no coldstarts, and affordable pricing.
Chatterbox Speech to Speech is a fast AI voice conversion model that converts source audio into a target voice style with optional reference audio guidance. Ready-to-use REST inference API for voice conversion, speech style transfer, dubbing, character voices, creator content, audio localization, and professional speech-to-speech workflows with simple integration, no coldstarts, and affordable pricing.
Chatterbox Text to Speech is a fast AI TTS model that converts text into expressive speech with optional reference audio, emotive tags, and delivery controls. Ready-to-use REST inference API for voice generation, narration, character dialogue, dubbing, virtual assistants, creator content, and professional text-to-speech workflows with simple integration, no coldstarts, and affordable pricing.
ByteDance Seed Speech TTS 2.0 is a fast AI text-to-speech model that converts text into natural speech with multilingual voices, delivery controls, and MP3 or Opus output. Ready-to-use REST inference API for voice generation, narration, dubbing, virtual assistants, product demos, creator content, and professional TTS workflows with simple integration, no coldstarts, and affordable pricing.
Kling Text-to-Audio turns text prompts into custom sound effects for videos, games, and multimedia using KlingAI's audio model. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Inworld Realtime TTS-2 converts text into low-latency, natural speech with official TTS-2 controls for delivery mode, language, timestamps, text normalization, and audio output settings. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Gemini 3.1 Flash Text to Speech generates expressive multi-speaker audio from text, with natural voices and multilingual language control for dialogue, narration, localization, and AI voice workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Minimax Speech 2.6 Turbo is a Text-to-Speech model offering ultra-human voice cloning, industry-leading text normalization, sub-250ms latency and 40+ language support. Pricing: $0.06 per 1000 characters. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
MiniMax Speech 2.8 Turbo is a high-definition text-to-speech model with natural and expressive voice synthesis. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ElevenLabs Turbo V2 is a Text-To-Speech model available via WaveSpeedAI, billed at $0.05 per 1000 characters for API requests. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ElevenLabs Flash V2 is a Text-to-Speech model that converts text into spoken audio using the ElevenLabs Flash V2 engine. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ElevenLabs Flash v2.5 is a text-to-speech model on WaveSpeedAI, billed at $0.05 per 1000 characters for generated speech. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Kling V1 TTS creates natural-sounding audio and supports KlingAI image, video, sound effect, virtual model, and custom AI workflows. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Qwen3 TTS Flash: Low-latency Text-to-Speech for English and Chinese with multiple voices, ideal for real-time dialogue. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ElevenLabs eleven-v3 is a text-to-speech model available as a hosted endpoint. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ACE-Step generates up to 4-minute music with lyrics from text and high acoustic fidelity; supports voice cloning, lyric edits, and remixing. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ACE-Step Audio-to-Audio turns existing tracks into remixes or vocal edits using remix and lyrics modes while preserving audio character. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ACE-Step Prompt-to-Audio creates music from simple prompts, auto-generating genre tags and lyrics for quick song creation. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ACE-Step Audio Outpaint generates seamless start or end extensions that match the original, ideal for intros, outros and longer tracks. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ACE-Step Audio Inpaint edits a specific audio segment to change lyrics or style while preserving the surrounding audio. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Minimax Speech 2.6 HD: Ultra-human, low-latency (< 250ms) TTS with voice cloning, text normalization and support for 40+ languages. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Minimax Speech 02 HD is Minimax's high-definition text-to-speech model delivering clear HD voices; pricing $0.05 per 1,000 characters. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ElevenLabs Turbo V2.5 is a text-to-speech model available via WaveSpeedAI, billed at $0.05 per 1000 characters for TTS requests. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
wavespeed-ai/vibevoice is an advanced voice generation model for producing high-fidelity, natural, and expressive speech from text, with optional speaker/region-style control for more precise results and easy integration into real-world applications. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
ElevenLabs Multilingual V2 is a multilingual text-to-speech model; cost $0.1 per 1000 characters. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
MiniMax Speech 2.8 HD is a high-definition text-to-speech model with natural and expressive voice synthesis for premium audio quality. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Qwen3 TTS: Multi-language, multi-voice text-to-speech synthesis with style control. Supports 11 languages and 9 voice characters. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Qwen3 TTS Voice Clone: Clone any voice from a reference audio and generate speech in that voice. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Qwen3 TTS Voice Design: Generate speech with custom voice characteristics described in natural language. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.
Microsoft VibeVoice text-to-speech model generates long-form speech from text with multi-speaker dialogue support. Choose from 9 voice presets across English, Chinese, and Hindi. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Inworld 1.5 Max delivers premium text-to-speech synthesis with 56+ multilingual voices, adjustable speaking rate, and high-fidelity natural-sounding audio output. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Inworld 1.5 Mini delivers high-quality text-to-speech synthesis with 56+ multilingual voices, adjustable speaking rate, and natural-sounding audio output. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Google Gemini 2.5 Pro Text-to-Speech delivers natural multi-speaker voice synthesis with 30+ voices across 24 languages. Perfect for dialogues, conversations, and multilingual content. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Google Gemini 2.5 Flash Text-to-Speech delivers fast, natural multi-speaker voice synthesis with 30+ voices across 24 languages at lower cost. Perfect for dialogues, conversations, and multilingual content. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
OmniVoice is a massively multilingual zero-shot TTS supporting 600+ languages. Generate speech with auto voice or design custom voices using natural language descriptions. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
OmniVoice Voice Clone clones any voice from a short 3-10 second audio sample. Supports 600+ languages with zero-shot voice cloning. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Minimax Speech 2.5 Turbo Preview: HD TTS with multilingual support, accurate voice replication across 40 languages. $0.04/1000 chars. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
MiniMax Speech 2.5 HD Preview offers HD TTS with enhanced multilingual expressiveness, accurate voice cloning, and 40-language support. Ready-to-use REST API, best performance, no coldstarts, affordable pricing.
MiniMax Voice Design generates natural voices from textual descriptions - no cloning - lets you set tone, accent and personality. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Minimax Voice Clone creates high-quality voice clones from short reference clips, closely matching tone, accent, and speaking style. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Minimax Speech-02 Turbo is a high-definition text-to-speech model delivering natural voice output. Cost: $0.03 per 1000 characters. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Запускайте любую модель из коллекции Speech Generation через один REST API. Оплата за генерацию — без подписок, без минимумов — с лидирующей в отрасли задержкой на инфраструктуре с аптаймом 99,9%.
Цена за вызов для каждой модели Speech Generation. Цена указана на странице каждой модели — без комиссий платформы сверху.
Большинство моделей изображений Speech Generation выполняются менее чем за 2 секунды. Видео- и 3D-модели в несколько раз быстрее самостоятельно размещаемых аналогов.
Мультирегиональная отказоустойчивость и автоматические повторы поддерживают ваш продакшен-трафик онлайн — даже при сбоях провайдера.
У каждой модели своя цена за вызов, указанная на странице модели. Мы выставляем счёт за каждую успешную генерацию, без подписок и минимумов.
Модели изображений в этой коллекции обычно выполняются менее чем за 2 секунды. Видео- и 3D-модели зависят от длительности и разрешения, но обычно в несколько раз быстрее самостоятельно размещаемых запусков.
Новые аккаунты, соответствующие условиям акции, могут получить 1 $ промокредитов для знакомства с моделями Speech Generation без банковской карты. Пробные кредиты не гарантируются при каждой регистрации; проверьте баланс перед генерацией.
Стандартные аккаунты имеют щедрые лимиты на параллельные задачи. Enterprise-планы предлагают индивидуальный RPM, более высокую конкурентность и выделенную мощность — свяжитесь с отделом продаж для подробностей.
Просмотрите наш полный каталог современных ИИ-моделей — изображения, видео, 3D, аудио, LLM и многое другое.
wavespeed.ai/models →Интегрируйте ИИ в свои приложения. RESTful API с клиентскими библиотеками — без холодных стартов, оплата по факту использования.
wavespeed.ai/docs →