Seedance 2.5 Now Live | Try in Video Generator →
openai
openai/gpt-5.6-luna

openai/gpt-5.6-luna

Release date: 2026-07-09

1,050,000 context · $0.20/M input tokens · $1.20/M output tokens

GPT-5.6 Luna is the lightweight model in OpenAI's GPT-5.6 series, designed for cost-efficient reasoning, coding, and agentic workflows at scale. It is well suited for high-throughput production workloads, lightweight automation, and large-volume application traffic where responsiveness and efficiency matter most.

Pricing

Pay-per-use

No upfront costs, pay only for what you use

Input
272K $0.20 / M Tokens
> 272K $0.40 / M Tokens
Output
272K $1.20 / M Tokens
> 272K $1.80 / M Tokens
Cache Read
272K $0.02 / M Tokens
> 272K $0.04 / M Tokens
Cache Write$0.25 / M Tokens

Try the model

openai/gpt-5.6-luna
Online
openai
Hi! I am a helpful AI assistant. What can I do for you?
Ready to use this model in a local coding agent?Agent setup

API Usage

Use the following code examples to integrate with our API:

import OpenAI from 'openai';

if (!process.env.WAVESPEED_API_KEY) throw new Error('Set WAVESPEED_API_KEY');
const client = new OpenAI({
  apiKey: process.env.WAVESPEED_API_KEY,
  baseURL: 'https://llm.wavespeed.ai/v1',
  timeout: 120_000,
  maxRetries: 2,
});

try {
  const response = await client.chat.completions.create({
    model: 'openai/gpt-5.6-luna',
    messages: [{ role: 'user', content: 'Hello!' }],
  });
  console.log(response.choices[0]?.message?.content ?? '');
} catch (error) {
  console.error('LLM request failed:', error);
  process.exitCode = 1;
}

Model Introduction

OpenAI: GPT-5.6 Luna

GPT-5.6 Luna is the lightweight model in OpenAI's GPT-5.6 series. It is designed for cost-efficient reasoning, coding, and agentic workflows where throughput and responsiveness matter more than using the highest-capability tier.

WaveSpeed AI exposes openai/gpt-5.6-luna through an OpenAI-compatible API, so it can be used with standard OpenAI SDKs and existing chat-completions-based application flows.


Why Use GPT-5.6 Luna

  • Lightweight GPT-5.6 model for high-volume workloads
  • Cost-efficient reasoning and coding for large-scale production use
  • Well suited for agentic workflows, tool use, and structured outputs
  • Supports long-context workloads such as document analysis and extended multi-turn sessions
  • Works with both Chat Completions and Responses API workflows

Key Features

  • Context Window: 1,050,000 tokens
  • Max Output: 128,000 tokens
  • Vision Input: Supported
  • Function Calling: Supported
  • Structured Outputs: Supported
  • Best Fit: cost-sensitive reasoning, coding, agents, high-throughput tasks

Specifications

SpecificationValue
ProviderOpenAI
Model IDopenai/gpt-5.6-luna
Model FamilyGPT-5.6
PositioningLightweight model
Context Window1,050,000 tokens
Max Output128,000 tokens
VisionSupported
Function CallingSupported
Structured OutputsSupported
Recommended Workloadscost-sensitive reasoning, coding, agentic workflows, high-throughput tasks

Pricing

Token TypeCost
Input$0.20 per million tokens
Cached Input$0.02 per million tokens
Cache Write$0.25 per million tokens
Output$1.20 per million tokens

Standard rates apply when the request contains no more than 272,000 input tokens.

Long-context pricing

Requests with 272,001 or more input tokens are billed at the following rates for the full request:

Token TypeCost
Input$0.40 per million tokens
Cached Input$0.04 per million tokens
Cache Write$0.50 per million tokens
Output$1.80 per million tokens

Cache writes are billed at 1.25× the applicable uncached input rate. Tool charges, including Web Search, are billed separately.


How to Use

Chat Completions

Use Chat Completions when you want a straightforward OpenAI-compatible integration path for standard conversational and coding workflows.

Python

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://llm.wavespeed.ai/v1"
)

response = client.chat.completions.create(
    model="openai/gpt-5.6-luna",
    messages=[
        {"role": "user", "content": "Summarize this issue in one paragraph."}
    ]
)

print(response.choices[0].message.content)

cURL

curl https://llm.wavespeed.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
    "model": "openai/gpt-5.6-luna",
    "messages": [
      {"role": "user", "content": "Summarize this issue in one paragraph."}
    ]
  }'

Pro Mode

GPT-5.6 Luna supports a stronger reasoning path through Pro mode.

Pro mode is not a separate core model that you need to configure independently. Instead, use the same base model, openai/gpt-5.6-luna, and enable Pro mode in the Responses API with:

{
  "reasoning": {
    "mode": "pro"
  }
}

Use Pro mode when you want the model to spend more effort on difficult reasoning, planning, and tool-using tasks. It is a better fit for complex coding, high-stakes decision logic, and multi-step agent workflows where answer quality matters more than speed or token efficiency.

In practice, Pro mode usually means:

  • higher reasoning depth
  • better consistency on difficult tasks
  • more token usage
  • higher latency than standard requests

Responses API Example

Python

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://llm.wavespeed.ai/v1"
)

response = client.responses.create(
    model="openai/gpt-5.6-luna",
    input="Review this automation design and identify the main reliability risk.",
    reasoning={
        "mode": "pro",
        "effort": "medium"
    }
)

print(response.output_text)

cURL

curl https://llm.wavespeed.ai/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
    "model": "openai/gpt-5.6-luna",
    "input": "Review this automation design and identify the main reliability risk.",
    "reasoning": {
      "mode": "pro",
      "effort": "medium"
    }
  }'

When to Use Pro Mode

Choose Pro mode for:

  • multi-step coding and debugging
  • agent workflows with tools or long chains of reasoning
  • tasks that require deeper analysis instead of quick turnaround
  • prompts where higher accuracy is worth additional cost and latency

Use standard mode when:

  • latency matters more than depth
  • the task is simple or repetitive
  • you are optimizing for throughput or cost

Notes

  • Designed as the most cost-efficient option within the GPT-5.6 family
  • Best paired with the Responses API when reasoning, tool use, or multi-turn state matters

Info

Provideropenai
Typellm

Supported Functionality

Input
TextImage
Output
Text
Context1,050,000
Max Output128,000
Vision✓ Supported
Function Calling✓ Supported

API Access Guide

Base URLhttps://llm.wavespeed.ai/v1
API Endpointchat/completions
Model IDopenai/gpt-5.6-luna

GPT 5.6 Luna API

openai/gpt-5.6-luna

GPT-5.6 Luna is the lightweight model in OpenAI's GPT-5.6 series, designed for cost-efficient reasoning, coding, and agentic workflows at scale. It is well suited for high-throughput production workloads, lightweight automation, and large-volume application traffic where responsiveness and efficiency matter most.

Input

$0.2 /M

Output

$1.2 /M

Context

1050K

Max Output

128K

Vision

Supported

Tool Use

Supported

Try GPT 5.6 Luna on WaveSpeedAI

Access GPT 5.6 Luna through our unified API — OpenAI-compatible, no cold starts, transparent pricing.

Frequently Asked Questions about GPT 5.6 Luna

How much does GPT 5.6 Luna cost via the API?+

Pricing on WaveSpeedAI: $0.20 per million input tokens and $1.20 per million output tokens. Prompt caching and batch processing are billed separately and reduce effective cost on long, repetitive workloads.

What is the context window of GPT 5.6 Luna?+

GPT 5.6 Luna supports up to 1050K tokens of context with up to 128K tokens of output per request.

Is GPT 5.6 Luna OpenAI-compatible?+

WaveSpeedAI exposes GPT 5.6 Luna at https://llm.wavespeed.ai/v1 through the OpenAI-compatible Chat Completions interface. Most OpenAI SDK clients work by changing the base URL and API key; optional fields depend on the selected model.

How do I get started with GPT 5.6 Luna?+

Sign in to WaveSpeedAI, create an API key in Access Keys, then send a request to https://llm.wavespeed.ai/v1/chat/completions with the model id shown above. Check the current model catalog for availability, capabilities, and pricing.

Related LLM APIs

GPT-5.6 Luna | OpenAI Frontier LLM API Pricing | WaveSpeedAI