WAN 3.0 is LIVE — 30s in one shot | Try in Video Generator →
openai
openai/gpt-5.6-terra

openai/gpt-5.6-terra

Release date: 2026-07-09

1,050,000 context · $2.00/M input tokens · $12.00/M output tokens

GPT-5.6 Terra is the balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is well suited for everyday coding, reasoning, and agentic workflows, offering a strong balance of quality, latency, and cost for general production use.

Pricing

Pay-per-use

No upfront costs, pay only for what you use

Input
272K $2.00 / M Tokens
> 272K $4.00 / M Tokens
Output
272K $12.00 / M Tokens
> 272K $18.00 / M Tokens
Cache Read
272K $0.20 / M Tokens
> 272K $0.40 / M Tokens
Cache Write$2.50 / M Tokens

Try the model

openai/gpt-5.6-terra
Online
openai
Hi! I am a helpful AI assistant. What can I do for you?
Ready to use this model in a local coding agent?Agent setup

API Usage

Use the following code examples to integrate with our API:

import OpenAI from 'openai';

if (!process.env.WAVESPEED_API_KEY) throw new Error('Set WAVESPEED_API_KEY');
const client = new OpenAI({
  apiKey: process.env.WAVESPEED_API_KEY,
  baseURL: 'https://llm.wavespeed.ai/v1',
  timeout: 120_000,
  maxRetries: 2,
});

try {
  const response = await client.chat.completions.create({
    model: 'openai/gpt-5.6-terra',
    messages: [{ role: 'user', content: 'Hello!' }],
  });
  console.log(response.choices[0]?.message?.content ?? '');
} catch (error) {
  console.error('LLM request failed:', error);
  process.exitCode = 1;
}

Model Introduction

OpenAI: GPT-5.6 Terra

GPT-5.6 Terra is the balanced model in OpenAI's GPT-5.6 series. It is designed for strong general-purpose reasoning, coding, and agentic workflows while offering a better balance of quality, latency, and cost than the flagship tier.

WaveSpeed AI exposes openai/gpt-5.6-terra through an OpenAI-compatible API, so it can be used with standard OpenAI SDKs and existing chat-completions-based application flows.


Why Use GPT-5.6 Terra

  • Balanced GPT-5.6 model for production workloads
  • Strong reasoning and coding performance with better cost efficiency than the flagship tier
  • Well suited for agentic workflows, tool use, and structured outputs
  • Supports long-context workloads such as document analysis and extended multi-turn sessions
  • Works with both Chat Completions and Responses API workflows

Key Features

  • Context Window: 1,050,000 tokens
  • Max Output: 128,000 tokens
  • Vision Input: Supported
  • Function Calling: Supported
  • Structured Outputs: Supported
  • Best Fit: general production reasoning, coding, agents, long-context tasks

Specifications

SpecificationValue
ProviderOpenAI
Model IDopenai/gpt-5.6-terra
Model FamilyGPT-5.6
PositioningBalanced model
Context Window1,050,000 tokens
Max Output128,000 tokens
VisionSupported
Function CallingSupported
Structured OutputsSupported
Recommended Workloadsgeneral reasoning, coding, agentic workflows, long-context tasks

Pricing

Token TypeCost
Input$2.00 per million tokens
Cached Input$0.20 per million tokens
Cache Write$2.50 per million tokens
Output$12 per million tokens

Standard rates apply when the request contains no more than 272,000 input tokens.

Long-context pricing

Requests with 272,001 or more input tokens are billed at the following rates for the full request:

Token TypeCost
Input$4.00 per million tokens
Cached Input$0.40 per million tokens
Cache Write$5.00 per million tokens
Output$18.00 per million tokens

Cache writes are billed at 1.25× the applicable uncached input rate. Tool charges, including Web Search, are billed separately.


How to Use

Chat Completions

Use Chat Completions when you want a straightforward OpenAI-compatible integration path for standard conversational and coding workflows.

Python

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://llm.wavespeed.ai/v1"
)

response = client.chat.completions.create(
    model="openai/gpt-5.6-terra",
    messages=[
        {"role": "user", "content": "Summarize this implementation plan in one paragraph."}
    ]
)

print(response.choices[0].message.content)

cURL

curl https://llm.wavespeed.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
    "model": "openai/gpt-5.6-terra",
    "messages": [
      {"role": "user", "content": "Summarize this implementation plan in one paragraph."}
    ]
  }'

Pro Mode

GPT-5.6 Terra supports a stronger reasoning path through Pro mode.

Pro mode is not a separate core model that you need to configure independently. Instead, use the same base model, openai/gpt-5.6-terra, and enable Pro mode in the Responses API with:

{
  "reasoning": {
    "mode": "pro"
  }
}

Use Pro mode when you want the model to spend more effort on difficult reasoning, planning, and tool-using tasks. It is a better fit for complex coding, high-stakes decision logic, and multi-step agent workflows where answer quality matters more than speed or token efficiency.

In practice, Pro mode usually means:

  • higher reasoning depth
  • better consistency on difficult tasks
  • more token usage
  • higher latency than standard requests

Responses API Example

Python

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://llm.wavespeed.ai/v1"
)

response = client.responses.create(
    model="openai/gpt-5.6-terra",
    input="Review this rollout plan and identify the main operational risk.",
    reasoning={
        "mode": "pro",
        "effort": "medium"
    }
)

print(response.output_text)

cURL

curl https://llm.wavespeed.ai/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
    "model": "openai/gpt-5.6-terra",
    "input": "Review this rollout plan and identify the main operational risk.",
    "reasoning": {
      "mode": "pro",
      "effort": "medium"
    }
  }'

When to Use Pro Mode

Choose Pro mode for:

  • multi-step coding and debugging
  • agent workflows with tools or long chains of reasoning
  • tasks that require deeper analysis instead of quick turnaround
  • prompts where higher accuracy is worth additional cost and latency

Use standard mode when:

  • latency matters more than depth
  • the task is simple or repetitive
  • you are optimizing for throughput or cost

Notes

  • Designed as a balanced option within the GPT-5.6 family
  • Best paired with the Responses API when reasoning, tool use, or multi-turn state matters

Info

Provideropenai
Typellm

Supported Functionality

Input
TextImage
Output
Text
Context1,050,000
Max Output128,000
Vision✓ Supported
Function Calling✓ Supported

API Access Guide

Base URLhttps://llm.wavespeed.ai/v1
API Endpointchat/completions
Model IDopenai/gpt-5.6-terra

GPT 5.6 Terra API

openai/gpt-5.6-terra

GPT-5.6 Terra is the balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is well suited for everyday coding, reasoning, and agentic workflows, offering a strong balance of quality, latency, and cost for general production use.

Input

$2 /M

Output

$12 /M

Context

1050K

Max Output

128K

Vision

Supported

Tool Use

Supported

Try GPT 5.6 Terra on WaveSpeedAI

Access GPT 5.6 Terra through our unified API — OpenAI-compatible, no cold starts, transparent pricing.

Frequently Asked Questions about GPT 5.6 Terra

How much does GPT 5.6 Terra cost via the API?+

Pricing on WaveSpeedAI: $2.00 per million input tokens and $12.00 per million output tokens. Prompt caching and batch processing are billed separately and reduce effective cost on long, repetitive workloads.

What is the context window of GPT 5.6 Terra?+

GPT 5.6 Terra supports up to 1050K tokens of context with up to 128K tokens of output per request.

Is GPT 5.6 Terra OpenAI-compatible?+

WaveSpeedAI exposes GPT 5.6 Terra at https://llm.wavespeed.ai/v1 through the OpenAI-compatible Chat Completions interface. Most OpenAI SDK clients work by changing the base URL and API key; optional fields depend on the selected model.

How do I get started with GPT 5.6 Terra?+

Sign in to WaveSpeedAI, create an API key in Access Keys, then send a request to https://llm.wavespeed.ai/v1/chat/completions with the model id shown above. Check the current model catalog for availability, capabilities, and pricing.

Related LLM APIs

GPT-5.6 Terra | OpenAI Frontier LLM API Pricing | WaveSpeedAI