Any Llm API Documentation
Playground
Try it on WaveSpeedAI!Any LLM is a versatile large language model for text generation, comprehension, and diverse NLP tasks such as chat and summarization. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.
Features
Any LLM is a unified text-generation endpoint. Use one request format to run a curated set of language models.
Supported models
- google/gemini-2.5-flash
- google/gemini-2.5-pro
- google/gemini-3-flash-preview
- qwen/qwen3.6-35b-a3b
- qwen/qwen3.6-27b
- qwen/qwen3.7-flash
Model selection and fallback
Pass one of the exact model IDs above in the model field. If model is omitted, the schema default is used. If a supplied model ID is not in the supported list, the request is automatically routed to qwen/qwen3.7-flash.
The supported model list and fallback policy for this endpoint may change at any time. Check the current model selector or API schema before relying on a specific model.
Parameters
| Parameter | Required | Description |
|---|---|---|
prompt | Yes | User prompt or instruction. |
system_prompt | No | System instructions, up to 10,000 characters. |
model | No | Exact supported model ID. Unknown IDs use the fallback above. |
reasoning | No | Include supported reasoning content in the final answer. |
priority | No | latency or throughput. |
temperature | No | Sampling temperature from 0 to 2. |
max_tokens | No | Maximum generated tokens, subject to the selected model context limit. |
enable_sync_mode | No | Attempt to wait for the result in the same API response. |
Notes
- Model capabilities, context limits, latency, and output behavior vary by provider.
- Processing time varies with the selected model and request complexity.
- Requests must comply with the applicable usage guidelines.
Authentication
For authentication details, please refer to the Authentication Guide.
API Endpoints
Submit Task & Query Result
set -euo pipefail
export WAVESPEED_API_KEY="your-api-key"
REQUEST_BODY=$(cat <<'JSON'
{
"prompt": "A cinematic ocean wave at sunrise, highly detailed",
"reasoning": false,
"priority": "latency",
"model": "google/gemini-2.5-flash"
}
JSON
)
# 1. Submit the prediction.
SUBMIT_RESPONSE=$(curl --silent --show-error --fail-with-body \
-X POST "https://api.wavespeed.ai/api/v3/wavespeed-ai/any-llm" \
-H "Authorization: Bearer ${WAVESPEED_API_KEY}" \
-H "Content-Type: application/json" \
-d "${REQUEST_BODY}")
TASK=$(printf '%s' "${SUBMIT_RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
PREDICTION_ID=$(printf '%s' "${TASK}" | jq -r '.id // empty')
if [ -z "${PREDICTION_ID}" ]; then
printf 'Submission response did not contain a prediction id
' >&2
exit 1
fi
RESULT_URL="https://api.wavespeed.ai/api/v3/predictions/${PREDICTION_ID}/result"
# 2. Poll until the prediction finishes.
while true; do
RESPONSE=$(curl --silent --show-error --fail-with-body \
"${RESULT_URL}" \
-H "Authorization: Bearer ${WAVESPEED_API_KEY}")
RESULT=$(printf '%s' "${RESPONSE}" | jq 'if type == "object" and has("data") then .data else . end')
STATUS=$(printf '%s' "${RESULT}" | jq -r '.status // empty')
case "${STATUS}" in
completed) printf '%s\n' "${RESULT}" | jq '.outputs'; break ;;
failed|cancelled|timeout|deleted) printf '%s\n' "${RESULT}" | jq . >&2; exit 1 ;;
*) sleep 2 ;;
esac
doneParameters
Task Submission Parameters
Request Parameters
| Parameter | Type | Required | Default | Range | Description |
|---|---|---|---|---|---|
| prompt | string | Yes | - | The positive prompt for the generation. | |
| system_prompt | string | No | - | - | System prompt to provide context or instructions to the model |
| reasoning | boolean | No | false | - | Should reasoning be the part of the final answer. |
| priority | string | No | latency | throughput, latency | Throughput is the default and is recommended for most use cases. Latency is recommended for use cases where low latency is important. |
| temperature | number | No | - | 0 ~ 2 | This setting influences the variety in the model's responses. Lower values lead to more predictable and typical responses, while higher values encourage more diverse and less common responses. At 0, the model always gives the same response for a given input. |
| max_tokens | integer | No | - | 1 ~ ∞ | This sets the upper limit for the number of tokens the model can generate in response. It won't produce more than this limit. The maximum value is the context length minus the prompt length. |
| model | string | No | google/gemini-2.5-flash | - | Model ID to use. Supported values are shown in the model selector. If a supplied model ID is not listed, the request falls back to qwen/qwen3.7-flash. The supported model list and fallback policy for this endpoint may change at any time. |
| enable_sync_mode | boolean | No | false | - | If set to `true`, the request attempts to wait for the generated result and return outputs in the same response. If the result is not ready within the sync wait window, the API can return a timeout body while the task continues processing. This option is only available via the API and is supported only by some models. |
Response Parameters
| Parameter | Type | Description |
|---|---|---|
| code | integer | HTTP status code (e.g., 200 for success) |
| message | string | Status message (e.g., “success”) |
| data.id | string | Unique identifier for the prediction, Task Id |
| data.model | string | Model ID used for the prediction |
| data.outputs | array | Output values, usually URL strings; some models return text strings or structured result objects (empty when status is not completed) |
| data.urls | object | Object containing related API endpoints |
| data.status | string | Task status. completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses. |
| data.created_at | string | ISO timestamp of when the request was created (e.g., “2023-04-01T12:34:56.789Z”) |
| data.error | string | Error message (empty if no error occurred) |
| data.timings | object | Object containing timing details |
| data.timings.inference | integer | Inference time in milliseconds |
Result Request Parameters
| Parameter | Type | Required | Default | Description |
|---|---|---|---|---|
| id | string | Yes | - | Task ID |
Result Response Parameters
| Parameter | Type | Description |
|---|---|---|
| code | integer | HTTP status code (e.g., 200 for success) |
| message | string | Status message (e.g., “success”) |
| data | object | The prediction data object containing all details |
| data.id | string | Unique identifier for the prediction |
| data.model | string | Model ID used for the prediction |
| data.outputs | array<string | object> | Array of generated outputs (empty when status is not completed). Items are usually URL strings, but may be text strings or structured result objects, depending on the model. |
| data.urls | object | Object containing related API endpoints |
| data.status | string | Status: completed is successful; failed, cancelled, timeout, and deleted are failure terminal statuses |
| data.created_at | string | ISO timestamp of when the request was created |
| data.error | string | Error message (empty if no error occurred) |
| data.timings | object | Object containing timing details |
| data.timings.inference | integer | Inference time in milliseconds |