Seedream 5.0 Pro ist LIVE | Jetzt im Bildgenerator testen →
moonshot
moonshotai/kimi-k2.6

moonshotai/kimi-k2.6

Veröffentlichungsdatum: 2026-04-20

262,144 context · $0.95/M input tokens · $4.00/M output tokens

Kimi K2.6 is Moonshot AI’s open-source native multimodal agentic model, designed for long-horizon coding, coding-driven UI/UX generation, proactive autonomous execution, and multi-agent orchestration. Built on a 1T-parameter Mixture-of-Experts architecture with 32B active parameters, it supports text and image inputs, a 262K-token context window, thinking mode, preserve-thinking workflows, function calling, and structured outputs. It is especially strong for complex end-to-end coding tasks across Python, Rust, Go, front-end engineering, DevOps, performance optimization, and agentic workflow automation.

Preise

Pay-per-Use

Keine Vorabkosten, zahlen Sie nur, was Sie nutzen

Eingabe$0.95 / M Tokens
Ausgabe$4.00 / M Tokens
Cache Read$0.16 / M Tokens

Modell ausprobieren

moonshotai/kimi-k2.6
Online
moonshot
Hallo! Ich bin ein hilfreicher KI-Assistent. Womit kann ich helfen?
Bereit, dieses Modell in einem lokalen Coding-Agent zu verwenden?Agent-Setup

API-Nutzung

Verwenden Sie die folgenden Codebeispiele zur Integration mit unserer API:

import OpenAI from 'openai';

if (!process.env.WAVESPEED_API_KEY) throw new Error('Set WAVESPEED_API_KEY');
const client = new OpenAI({
  apiKey: process.env.WAVESPEED_API_KEY,
  baseURL: 'https://llm.wavespeed.ai/v1',
  timeout: 120_000,
  maxRetries: 2,
});

try {
  const response = await client.chat.completions.create({
    model: 'moonshotai/kimi-k2.6',
    messages: [{ role: 'user', content: 'Hello!' }],
  });
  console.log(response.choices[0]?.message?.content ?? '');
} catch (error) {
  console.error('LLM request failed:', error);
  process.exitCode = 1;
}

Modelleinführung

MoonshotAI: Kimi K2.6

Kimi K2.6 is Moonshot AI’s open-source native multimodal agentic model, designed for long-horizon coding, coding-driven UI/UX generation, proactive autonomous execution, and multi-agent orchestration. Built on a 1T-parameter Mixture-of-Experts architecture with 32B active parameters, it is optimized for complex coding, visual understanding, tool use, and large-scale agent workflows.


Why It Looks Great

  • Open-source native multimodal agentic model from Moonshot AI
  • 1T-parameter Mixture-of-Experts architecture with 32B active parameters
  • 262K-token context window for long prompts, large codebases, documents, and multi-turn workflows
  • Strong long-horizon coding performance across Python, Rust, Go, front-end, DevOps, and optimization tasks
  • Excellent fit for coding-driven UI/UX generation, including full-stack apps and polished interfaces
  • Agent Swarm capabilities for decomposing and coordinating complex multi-agent workflows
  • Vision input support for screenshots, mockups, diagrams, and multimodal document understanding
  • Thinking mode and preserve-thinking support for multi-step reasoning and coding agent scenarios
  • Function calling and tool-use support for agentic application workflows
  • Structured output support for JSON responses and schema-constrained generation

Key Features

  • Architecture: Mixture-of-Experts
  • Total Parameters: 1T
  • Active Parameters: 32B
  • Context Window: 262,144 tokens
  • Max Input: Not listed
  • Max Output: Not listed
  • Input: Text, Image
  • Output: Text
  • Vision: Supported
  • Function Calling: Supported
  • Structured Outputs: Supported
  • Thinking Mode: Supported
  • Preserve Thinking: Supported
  • Image Generation: Not listed
  • Audio Input: Not listed
  • Supported Parameters: frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, parallel_tool_calls, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p

Specifications

SpecificationValue
Providermoonshot
Model TypeChat Completions model
ArchitectureMixture-of-Experts
Parameters1T total / 32B active
Experts384 experts, 8 selected per token
AttentionMLA
Vision EncoderMoonViT
Context Window262,144 tokens
InputText, Image
OutputText
VisionSupported
Function CallingSupported
Structured OutputsSupported
Thinking ModeSupported

Pricing

Token TypeCost
Input$0.73 per million tokens
Output$3.49 per million tokens
Cached Input$0.25 per million tokens

How to Use

  1. Write your prompt - describe the task, provide context, and specify the desired output format.
  2. Submit - the model processes your request and returns the response.

API Integration

Base URL: https://llm.wavespeed.ai/v1
API Endpoint: chat/completions
Model ID: moonshotai/kimi-k2.6


API Usage

Python SDK

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://llm.wavespeed.ai/v1"
)

response = client.chat.completions.create(
    model="moonshotai/kimi-k2.6",
    messages=[{"role": "user", "content": "Hello!"}]
)

print(response.choices[0].message.content)

cURL

curl https://llm.wavespeed.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
    "model": "moonshotai/kimi-k2.6",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Notes

  • Model: moonshotai/kimi-k2.6
  • Provider: moonshot
  • Best suited for long-horizon coding, UI/UX generation, visual understanding, tool use, multi-agent orchestration, and autonomous workflow execution

Info

Anbietermoonshot
Typllm

Unterstützte Funktionen

Eingabe
TextBild
Ausgabe
Text
Kontext262,144
Max. Ausgabe262,142
Vision✓ Unterstützt
Function Calling✓ Unterstützt

API-Zugriffsanleitung

Base URLhttps://llm.wavespeed.ai/v1
API-Endpunktchat/completions
Modell-IDmoonshotai/kimi-k2.6

Kimi K2.6 API

moonshotai/kimi-k2.6

Kimi K2.6 is Moonshot AI’s open-source native multimodal agentic model, designed for long-horizon coding, coding-driven UI/UX generation, proactive autonomous execution, and multi-agent orchestration. Built on a 1T-parameter Mixture-of-Experts architecture with 32B active parameters, it supports text and image inputs, a 262K-token context window, thinking mode, preserve-thinking workflows, function calling, and structured outputs. It is especially strong for complex end-to-end coding tasks across Python, Rust, Go, front-end engineering, DevOps, performance optimization, and agentic workflow automation.

Eingabe

$0.95 /M

Ausgabe

$4 /M

Kontext

262K

Max. Ausgabe

262K

Vision

Unterstützt

Tool-Nutzung

Unterstützt

Kimi K2.6 auf WaveSpeedAI testen

Zugriff auf Kimi K2.6 über unsere einheitliche API — OpenAI-kompatibel, keine Kaltstarts, transparente Preise.

Häufige Fragen zu Kimi K2.6

Wie viel kostet die Kimi K2.6-API?+

Preise auf WaveSpeedAI: $0.95 pro Million Input-Tokens und $4.00 pro Million Output-Tokens. Prompt-Caching und Batch-Verarbeitung werden separat berechnet und reduzieren die effektiven Kosten bei langen, sich wiederholenden Workloads.

Wie groß ist das Kontextfenster von Kimi K2.6?+

Kimi K2.6 unterstützt bis zu 262K Kontext-Tokens und bis zu 262K Output-Tokens pro Anfrage.

Ist Kimi K2.6 OpenAI-kompatibel?+

WaveSpeedAI stellt Kimi K2.6 unter https://llm.wavespeed.ai/v1 über die OpenAI-kompatible Chat-Completions-Schnittstelle bereit. Bei den meisten OpenAI-SDK-Clients reichen Base-URL und API-Schlüssel; optionale Felder hängen vom Modell ab.

Wie starte ich mit Kimi K2.6?+

Melden Sie sich bei WaveSpeedAI an, erstellen Sie unter Access Keys einen API-Schlüssel und senden Sie eine Anfrage mit der oben gezeigten Modell-ID an https://llm.wavespeed.ai/v1/chat/completions. Verfügbarkeit, Fähigkeiten und Preise finden Sie im aktuellen Modellkatalog.

Verwandte LLM-APIs

MoonshotAI: Kimi K2.6 | Moonshot Multimodal LLM API Pricing | WaveSpeedAI