Seedream 5.0 Pro is LIVE | Try in Image Generator →
moonshot
moonshotai/kimi-k3

moonshotai/kimi-k3

Release date: 2026-07-16

1,048,576 context · $3.00/M input tokens · $15.00/M output tokens

Kimi K3 is Moonshot AI's flagship open-weight multimodal reasoning model. It is built for complex coding, knowledge work, and long-horizon agentic workflows, with strong performance on large codebase navigation, tool use, debugging, visual reasoning, and iterative problem solving. WaveSpeed AI exposes moonshotai/kimi-k3 through an OpenAI-compatible API, so it can be used with standard OpenAI SDKs and existing chat-completions-based application flows.

Pricing

Pay-per-use

No upfront costs, pay only for what you use

Input$3.00 / M Tokens
Output$15.00 / M Tokens
Cache Read$0.30 / M Tokens

Try the model

moonshotai/kimi-k3
Online
moonshot
Hi! I am a helpful AI assistant. What can I do for you?
Ready to use this model in a local coding agent?Agent setup

API Usage

Use the following code examples to integrate with our API:

import OpenAI from 'openai';

if (!process.env.WAVESPEED_API_KEY) throw new Error('Set WAVESPEED_API_KEY');
const client = new OpenAI({
  apiKey: process.env.WAVESPEED_API_KEY,
  baseURL: 'https://llm.wavespeed.ai/v1',
  timeout: 120_000,
  maxRetries: 2,
});

try {
  const response = await client.chat.completions.create({
    model: 'moonshotai/kimi-k3',
    messages: [{ role: 'user', content: 'Hello!' }],
  });
  console.log(response.choices[0]?.message?.content ?? '');
} catch (error) {
  console.error('LLM request failed:', error);
  process.exitCode = 1;
}

Model Introduction

Moonshot AI: Kimi K3

Kimi K3 is Moonshot AI's flagship open-weight multimodal reasoning model. It is designed for complex coding, knowledge work, and long-horizon agentic workflows, and is especially strong at large repository understanding, tool use, debugging, and iterative work across images, logs, tests, and runtime feedback.

WaveSpeed AI exposes moonshotai/kimi-k3 through an OpenAI-compatible API, so it can be used with standard OpenAI SDKs and existing chat-completions-based application flows.


Why Use Kimi K3

  • Flagship Kimi model for advanced reasoning and software work
  • Strong long-horizon coding performance across large repositories and multi-step tasks
  • Well suited for agentic workflows, tool use, and structured outputs
  • Native multimodal capability for image-based understanding and visual iteration
  • Long-context support for document analysis, code review, and extended multi-turn sessions

Key Features

  • Context Window: 1,000,000 tokens
  • Max Output: up to 131,072 tokens by default
  • Vision Input: Supported
  • Function Calling: Supported
  • Structured Outputs: Supported
  • Reasoning: Enabled by default
  • Best Fit: coding, reasoning, agents, multimodal workflows, long-context tasks

Specifications

SpecificationValue
ProviderMoonshot AI
Model IDmoonshotai/kimi-k3
Model FamilyKimi K3
PositioningFlagship open-weight multimodal reasoning model
Parameters2.8T
Context Window1,000,000 tokens
Max Output131,072 tokens by default
VisionSupported
Function CallingSupported
Structured OutputsSupported
Recommended Workloadscomplex coding, reasoning, agentic workflows, multimodal analysis, long-context tasks

Architecture Notes

Kimi K3 is built with KDA (Kimi Delta Attention) and Attention Residuals to improve computational efficiency at scale. It is positioned for demanding workflows such as long-horizon programming, knowledge-intensive tasks, and multimodal reasoning over both text and images.


How to Use

Chat Completions

Use Chat Completions when you want a straightforward OpenAI-compatible integration path for conversational, coding, and agent workflows.

Python

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://llm.wavespeed.ai/v1"
)

response = client.chat.completions.create(
    model="moonshotai/kimi-k3",
    messages=[
        {"role": "user", "content": "Review this bug report and identify the most likely root cause."}
    ]
)

print(response.choices[0].message.content)

cURL

curl https://llm.wavespeed.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
    "model": "moonshotai/kimi-k3",
    "messages": [
      {"role": "user", "content": "Review this bug report and identify the most likely root cause."}
    ]
  }'

Multimodal Example

Kimi K3 supports native visual understanding, making it a strong fit for tasks such as screenshot debugging, UI review, diagram analysis, and image-grounded reasoning.

Python

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://llm.wavespeed.ai/v1"
)

response = client.chat.completions.create(
    model="moonshotai/kimi-k3",
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Describe the issue shown in this screenshot."},
                {
                    "type": "image_url",
                    "image_url": {
                        "url": "data:image/png;base64,BASE64_IMAGE_DATA"
                    }
                }
            ]
        }
    ]
)

print(response.choices[0].message.content)

Tool Use and Structured Output

Kimi K3 is well suited for applications that combine reasoning with tools and schema-constrained outputs.

Common use cases include:

  • debugging agents that inspect logs, test output, and code
  • repository assistants that navigate large codebases
  • multimodal workflows that combine screenshots with implementation tasks
  • structured extraction pipelines that require JSON output

Notes

  • Official upstream model name is kimi-k3
  • WaveSpeed model ID is drafted here as moonshotai/kimi-k3
  • Reasoning is part of the model's default behavior
  • Best paired with agentic and long-context workflows where tool use and iterative refinement matter

Info

Providermoonshot
Typellm

Supported Functionality

Input
TextImage
Output
Text
Context1,048,576
Max Output-
Vision✓ Supported
Function Calling✓ Supported

API Access Guide

Base URLhttps://llm.wavespeed.ai/v1
API Endpointchat/completions
Model IDmoonshotai/kimi-k3

Kimi K3 API

moonshotai/kimi-k3

Kimi K3 is Moonshot AI's flagship open-weight multimodal reasoning model. It is built for complex coding, knowledge work, and long-horizon agentic workflows, with strong performance on large codebase navigation, tool use, debugging, visual reasoning, and iterative problem solving. WaveSpeed AI exposes `moonshotai/kimi-k3` through an OpenAI-compatible API, so it can be used with standard OpenAI SDKs and existing chat-completions-based application flows.

Input

$3 /M

Output

$15 /M

Context

1049K

Vision

Supported

Tool Use

Supported

Try Kimi K3 on WaveSpeedAI

Access Kimi K3 through our unified API — OpenAI-compatible, no cold starts, transparent pricing.

Frequently Asked Questions about Kimi K3

How much does Kimi K3 cost via the API?+

Pricing on WaveSpeedAI: $3.00 per million input tokens and $15.00 per million output tokens. Prompt caching and batch processing are billed separately and reduce effective cost on long, repetitive workloads.

What is the context window of Kimi K3?+

Kimi K3 supports up to 1049K tokens of context with up to — tokens of output per request.

Is Kimi K3 OpenAI-compatible?+

WaveSpeedAI exposes Kimi K3 at https://llm.wavespeed.ai/v1 through the OpenAI-compatible Chat Completions interface. Most OpenAI SDK clients work by changing the base URL and API key; optional fields depend on the selected model.

How do I get started with Kimi K3?+

Sign in to WaveSpeedAI, create an API key in Access Keys, then send a request to https://llm.wavespeed.ai/v1/chat/completions with the model id shown above. Check the current model catalog for availability, capabilities, and pricing.

Related LLM APIs

MoonshotAI: Kimi K3 | Moonshot Multimodal LLM API Pricing | WaveSpeedAI