openai/gpt-5.6-luna
Release date: 2026-07-09
1,050,000 context · $0.20/M input tokens · $1.20/M output tokens
GPT-5.6 Luna is the lightweight model in OpenAI's GPT-5.6 series, designed for cost-efficient reasoning, coding, and agentic workflows at scale. It is well suited for high-throughput production workloads, lightweight automation, and large-volume application traffic where responsiveness and efficiency matter most.
Pay-per-use
No upfront costs, pay only for what you use
Use the following code examples to integrate with our API:
import OpenAI from 'openai';
if (!process.env.WAVESPEED_API_KEY) throw new Error('Set WAVESPEED_API_KEY');
const client = new OpenAI({
apiKey: process.env.WAVESPEED_API_KEY,
baseURL: 'https://llm.wavespeed.ai/v1',
timeout: 120_000,
maxRetries: 2,
});
try {
const response = await client.chat.completions.create({
model: 'openai/gpt-5.6-luna',
messages: [{ role: 'user', content: 'Hello!' }],
});
console.log(response.choices[0]?.message?.content ?? '');
} catch (error) {
console.error('LLM request failed:', error);
process.exitCode = 1;
}import OpenAI from 'openai';
if (!process.env.WAVESPEED_API_KEY) throw new Error('Set WAVESPEED_API_KEY');
const client = new OpenAI({
apiKey: process.env.WAVESPEED_API_KEY,
baseURL: 'https://llm.wavespeed.ai/v1',
timeout: 120_000,
maxRetries: 2,
});
try {
const response = await client.chat.completions.create({
model: 'openai/gpt-5.6-luna',
messages: [{ role: 'user', content: 'Hello!' }],
});
console.log(response.choices[0]?.message?.content ?? '');
} catch (error) {
console.error('LLM request failed:', error);
process.exitCode = 1;
}GPT-5.6 Luna is the lightweight model in OpenAI's GPT-5.6 series. It is designed for cost-efficient reasoning, coding, and agentic workflows where throughput and responsiveness matter more than using the highest-capability tier.
WaveSpeed AI exposes openai/gpt-5.6-luna through an OpenAI-compatible API, so it can be used with standard OpenAI SDKs and existing chat-completions-based application flows.
| Specification | Value |
|---|---|
| Provider | OpenAI |
| Model ID | openai/gpt-5.6-luna |
| Model Family | GPT-5.6 |
| Positioning | Lightweight model |
| Context Window | 1,050,000 tokens |
| Max Output | 128,000 tokens |
| Vision | Supported |
| Function Calling | Supported |
| Structured Outputs | Supported |
| Recommended Workloads | cost-sensitive reasoning, coding, agentic workflows, high-throughput tasks |
| Token Type | Cost |
|---|---|
| Input | $0.20 per million tokens |
| Cached Input | $0.02 per million tokens |
| Cache Write | $0.25 per million tokens |
| Output | $1.20 per million tokens |
Standard rates apply when the request contains no more than 272,000 input tokens.
Requests with 272,001 or more input tokens are billed at the following rates for the full request:
| Token Type | Cost |
|---|---|
| Input | $0.40 per million tokens |
| Cached Input | $0.04 per million tokens |
| Cache Write | $0.50 per million tokens |
| Output | $1.80 per million tokens |
Cache writes are billed at 1.25× the applicable uncached input rate. Tool charges, including Web Search, are billed separately.
Use Chat Completions when you want a straightforward OpenAI-compatible integration path for standard conversational and coding workflows.
Python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://llm.wavespeed.ai/v1"
)
response = client.chat.completions.create(
model="openai/gpt-5.6-luna",
messages=[
{"role": "user", "content": "Summarize this issue in one paragraph."}
]
)
print(response.choices[0].message.content)
cURL
curl https://llm.wavespeed.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "openai/gpt-5.6-luna",
"messages": [
{"role": "user", "content": "Summarize this issue in one paragraph."}
]
}'
GPT-5.6 Luna supports a stronger reasoning path through Pro mode.
Pro mode is not a separate core model that you need to configure independently. Instead, use the same base model, openai/gpt-5.6-luna, and enable Pro mode in the Responses API with:
{
"reasoning": {
"mode": "pro"
}
}
Use Pro mode when you want the model to spend more effort on difficult reasoning, planning, and tool-using tasks. It is a better fit for complex coding, high-stakes decision logic, and multi-step agent workflows where answer quality matters more than speed or token efficiency.
In practice, Pro mode usually means:
Python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://llm.wavespeed.ai/v1"
)
response = client.responses.create(
model="openai/gpt-5.6-luna",
input="Review this automation design and identify the main reliability risk.",
reasoning={
"mode": "pro",
"effort": "medium"
}
)
print(response.output_text)
cURL
curl https://llm.wavespeed.ai/v1/responses \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"model": "openai/gpt-5.6-luna",
"input": "Review this automation design and identify the main reliability risk.",
"reasoning": {
"mode": "pro",
"effort": "medium"
}
}'
Choose Pro mode for:
Use standard mode when:
openai/gpt-5.6-luna
GPT-5.6 Luna is the lightweight model in OpenAI's GPT-5.6 series, designed for cost-efficient reasoning, coding, and agentic workflows at scale. It is well suited for high-throughput production workloads, lightweight automation, and large-volume application traffic where responsiveness and efficiency matter most.
Input
$0.2 /M
Output
$1.2 /M
Context
1050K
Max Output
128K
Vision
Supported
Tool Use
Supported
Access GPT 5.6 Luna through our unified API — OpenAI-compatible, no cold starts, transparent pricing.
Pricing on WaveSpeedAI: $0.20 per million input tokens and $1.20 per million output tokens. Prompt caching and batch processing are billed separately and reduce effective cost on long, repetitive workloads.
GPT 5.6 Luna supports up to 1050K tokens of context with up to 128K tokens of output per request.
WaveSpeedAI exposes GPT 5.6 Luna at https://llm.wavespeed.ai/v1 through the OpenAI-compatible Chat Completions interface. Most OpenAI SDK clients work by changing the base URL and API key; optional fields depend on the selected model.
Sign in to WaveSpeedAI, create an API key in Access Keys, then send a request to https://llm.wavespeed.ai/v1/chat/completions with the model id shown above. Check the current model catalog for availability, capabilities, and pricing.