# minimax/speech-2.6-turbo

> Minimax Speech 2.6 Turbo is a Text-to-Speech model offering ultra-human voice cloning, industry-leading text normalization, sub-250ms latency and 40+ language support. Pricing: $0.06 per 1000 characters. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

## Overview

- **Endpoint**: `https://api.wavespeed.ai/api/v3/minimax/speech-2.6-turbo`
- **Polling/result URL**: `https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result`
- **Model ID**: `minimax/speech-2.6-turbo`
- **Category**: text-to-speech

## API Information

This model can be used via our HTTP API or more conveniently via our client libraries.
The API is asynchronous: submit a prediction, then poll its result URL until it completes.

### Input Schema

The API accepts the following input parameters:

- **`text`** (`string`, _required_):
  Text to convert to speech. Every character is 1 token. Maximum 10000 characters. Use <#x#> between words to control pause duration (0.01-99.99s).

- **`voice_id`** (`string`, _required_):
  Desired voice ID. Use a voice ID you have trained (https://wavespeed.ai/models/minimax/voice-clone), or one of the following system voice IDs: Wise_Woman, Friendly_Person, Inspirational_girl, Deep_Voice_Man, Calm_Woman, Casual_Guy, Lively_Girl, Patient_Man, Young_Knight, Determined_Man, Lovely_Girl, Decent_Boy, Imposing_Manner, Elegant_Man, Abbess, Sweet_Girl_2, Exuberant_Girl.

- **`speed`** (`number`, _optional_):
  Speech speed. Range: 0.5-2.0, where 1.0 is normal speed.
  - Default: `1`
  - Range: `0.5` to `2`

- **`volume`** (`number`, _optional_):
  Speech volume. Range: 0.1-10.0, where 1.0 is normal volume.
  - Default: `1`
  - Range: `0.1` to `10`

- **`pitch`** (`number`, _optional_):
  Speech pitch. Range: -12 to 12, where 0 is normal pitch.
  - Default: `0`
  - Range: `-12` to `12`

- **`emotion`** (`string`, _optional_):
  The emotion of the generated speech.
  - Default: `"happy"`
  - Options: "happy", "sad", "angry", "fearful", "disgusted", "surprised", "neutral"

- **`english_normalization`** (`boolean`, _optional_):
  This parameter supports English text normalization, which improves performance in number-reading scenarios.
  - Default: `false`

- **`sample_rate`** (`integer`, _optional_):
  Sample rate of generated sound.
  - Options: 8000, 16000, 22050, 24000, 32000, 44100

- **`bitrate`** (`integer`, _optional_):
  Bitrate of generated sound.
  - Options: 32000, 64000, 128000, 256000

- **`channel`** (`string`, _optional_):
  The number of channels of the generated audio. 1: mono, 2: stereo.
  - Options: "1", "2"

- **`format`** (`string`, _optional_):
  Format of generated sound.
  - Options: "mp3", "wav", "pcm", "flac"

- **`language_boost`** (`string`, _optional_):
  Enhance the ability to recognize specified languages and dialects.
  - Options: "Chinese", "Chinese,Yue", "English", "Arabic", "Russian", "Spanish", "French", "Portuguese", "German", "Turkish", "Dutch", "Ukrainian", "Vietnamese", "Indonesian", "Japanese", "Italian", "Korean", "Thai", "Polish", "Romanian", "Greek", "Czech", "Finnish", "Hindi", "Bulgarian", "Danish", "Hebrew", "Malay", "Persian", "Slovak", "Swedish", "Croatian", "Filipino", "Hungarian", "Norwegian", "Slovenian", "Catalan", "Nynorsk", "Tamil", "Afrikaans", "auto"

- **`enable_base64_output`** (`boolean`, _optional_):
  If set to `true`, the prediction's `output` strings are returned as **naked base64** (no `data:<mime>;base64,` prefix). When `false` (default), outputs are returned as URLs pointing to our CDN.
  - Default: `false`

- **`enable_sync_mode`** (`boolean`, _optional_):
  If set to `true`, the request attempts to wait for the generated result and return outputs in the same response. If the result is not ready within the sync wait window, the API can return a timeout body while the task continues processing. This option is only available via the API and is supported only by some models.
  - Default: `false`



**Required Parameters Example**:

```json
{
  "text": "A clear example input",
  "voice_id": "example"
}
```

**Full Example**:

```json
{
  "text": "A clear example input",
  "voice_id": "example",
  "speed": 1,
  "volume": 1,
  "pitch": 0,
  "emotion": "happy",
  "english_normalization": false,
  "sample_rate": 8000,
  "bitrate": 32000,
  "channel": "1",
  "format": "mp3",
  "language_boost": "Chinese"
}
```

### Result Data Schema

The `data` object returned by the API has the following fields:

- **`created_at`** (`string (date-time)`, _optional_):
  ISO timestamp of when the request was created (e.g., "2023-04-01T12:34:56.789Z").

- **`id`** (`string`, _optional_):
  Unique identifier for the prediction, the ID of the prediction to get.

- **`model`** (`string`, _optional_):
  Model ID used for the prediction.

- **`outputs`** (`array of string | object`, _optional_):
  Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model.

- **`status`** (`string`, _optional_):
  Status of the task: created, processing, completed, or failed.

- **`urls`** (`object`, _optional_):
  Object containing related API endpoints.



**Example `data` Object**:

```json
{
  "created_at": "example",
  "id": "example",
  "model": "example",
  "outputs": [],
  "status": "example",
  "urls": {}
}
```

## Usage Examples

The examples use `jq` to read JSON. Set your API key first:

```bash
set -euo pipefail
export WAVESPEED_API_KEY="your-api-key"
```

### 1. Submit a prediction

```bash
REQUEST_BODY=$(cat <<'JSON'
{
  "text": "A clear example input",
  "voice_id": "example"
}
JSON
)

SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
  --request POST \
  --url https://api.wavespeed.ai/api/v3/minimax/speech-2.6-turbo \
  --header "Authorization: Bearer ${WAVESPEED_API_KEY}" \
  --header "Content-Type: application/json" \
  --data "${REQUEST_BODY}")

printf '%s\n' "${SUBMIT_RESPONSE}" | jq .
```

The response contains the prediction ID in `data.id`.

### 2. Poll until complete and read `outputs`

```bash
PREDICTION_ID=$(printf '%s' "${SUBMIT_RESPONSE}" | jq -r '.data.id')
if [ -z "${PREDICTION_ID}" ] || [ "${PREDICTION_ID}" = "null" ]; then
  printf 'Submission response did not contain data.id\n' >&2
  exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"

while true; do
  RESPONSE=$(curl --silent --show-error --fail-with-body \
    --request GET \
    --url "${RESULT_URL}" \
    --header "Authorization: Bearer ${WAVESPEED_API_KEY}")

  RESULT=$(printf '%s' "${RESPONSE}" | jq -e '.data')
  STATUS=$(printf '%s' "${RESULT}" | jq -er '.status')
  case "${STATUS}" in
    completed)
      # Generated files are returned in the outputs array.
      printf '%s\n' "${RESULT}" | jq '.outputs'
      break
      ;;
    failed|cancelled|timeout|deleted)
      printf '%s\n' "${RESULT}" | jq '{status, error, code}'
      exit 1
      ;;
    *)
      sleep 2
      ;;
  esac
done
```

## Additional Resources

### Documentation

- [Model Playground](https://wavespeed.ai/models/minimax/speech-2.6-turbo)
- [API Documentation](https://wavespeed.ai/docs/docs-api/minimax/minimax-speech-2.6-turbo)
