Seedance 2.0 15% DE DESCONTO | Crie no Video Generator →
moonshot
moonshotai/kimi-k2.6

moonshotai/kimi-k2.6

Data de lançamento: 2026-04-20

262,144 context · $0.95/M input tokens · $4.00/M output tokens

Kimi K2.6 is Moonshot AI’s open-source native multimodal agentic model, designed for long-horizon coding, coding-driven UI/UX generation, proactive autonomous execution, and multi-agent orchestration. Built on a 1T-parameter Mixture-of-Experts architecture with 32B active parameters, it supports text and image inputs, a 262K-token context window, thinking mode, preserve-thinking workflows, function calling, and structured outputs. It is especially strong for complex end-to-end coding tasks across Python, Rust, Go, front-end engineering, DevOps, performance optimization, and agentic workflow automation.

Preços

Pagamento por uso

Sem custo inicial, pague apenas pelo que usar

Entrada$0.95 / M Tokens
Saída$4.00 / M Tokens
Cache Read$0.16 / M Tokens

Experimentar o modelo

moonshotai/kimi-k2.6
Online
moonshot
Olá! Sou um assistente de IA útil. Em que posso ajudar?

Uso da API

Use os exemplos de código abaixo para integrar com nossa API:

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://llm.wavespeed.ai/v1"
)

response = client.chat.completions.create(
    model="moonshotai/kimi-k2.6",
    messages=[
        {"role": "user", "content": "Hello!"}
    ]
)

print(response.choices[0].message.content)

Introdução do modelo

MoonshotAI: Kimi K2.6

Kimi K2.6 is Moonshot AI’s open-source native multimodal agentic model, designed for long-horizon coding, coding-driven UI/UX generation, proactive autonomous execution, and multi-agent orchestration. Built on a 1T-parameter Mixture-of-Experts architecture with 32B active parameters, it is optimized for complex coding, visual understanding, tool use, and large-scale agent workflows.


Why It Looks Great

  • Open-source native multimodal agentic model from Moonshot AI
  • 1T-parameter Mixture-of-Experts architecture with 32B active parameters
  • 262K-token context window for long prompts, large codebases, documents, and multi-turn workflows
  • Strong long-horizon coding performance across Python, Rust, Go, front-end, DevOps, and optimization tasks
  • Excellent fit for coding-driven UI/UX generation, including full-stack apps and polished interfaces
  • Agent Swarm capabilities for decomposing and coordinating complex multi-agent workflows
  • Vision input support for screenshots, mockups, diagrams, and multimodal document understanding
  • Thinking mode and preserve-thinking support for multi-step reasoning and coding agent scenarios
  • Function calling and tool-use support for agentic application workflows
  • Structured output support for JSON responses and schema-constrained generation

Key Features

  • Architecture: Mixture-of-Experts
  • Total Parameters: 1T
  • Active Parameters: 32B
  • Context Window: 262,144 tokens
  • Max Input: Not listed
  • Max Output: Not listed
  • Input: Text, Image
  • Output: Text
  • Vision: Supported
  • Function Calling: Supported
  • Structured Outputs: Supported
  • Thinking Mode: Supported
  • Preserve Thinking: Supported
  • Image Generation: Not listed
  • Audio Input: Not listed
  • Supported Parameters: frequency_penalty, include_reasoning, logit_bias, logprobs, max_tokens, min_p, parallel_tool_calls, presence_penalty, reasoning, reasoning_effort, repetition_penalty, response_format, seed, stop, structured_outputs, temperature, tool_choice, tools, top_k, top_logprobs, top_p

Specifications

SpecificationValue
Providermoonshot
Model TypeChat Completions model
ArchitectureMixture-of-Experts
Parameters1T total / 32B active
Experts384 experts, 8 selected per token
AttentionMLA
Vision EncoderMoonViT
Context Window262,144 tokens
InputText, Image
OutputText
VisionSupported
Function CallingSupported
Structured OutputsSupported
Thinking ModeSupported

Pricing

Token TypeCost
Input$0.73 per million tokens
Output$3.49 per million tokens
Cached Input$0.25 per million tokens

How to Use

  1. Write your prompt - describe the task, provide context, and specify the desired output format.
  2. Submit - the model processes your request and returns the response.

API Integration

Base URL: https://llm.wavespeed.ai/v1
API Endpoint: chat/completions
Model ID: moonshotai/kimi-k2.6


API Usage

Python SDK

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://llm.wavespeed.ai/v1"
)

response = client.chat.completions.create(
    model="moonshotai/kimi-k2.6",
    messages=[{"role": "user", "content": "Hello!"}]
)

print(response.choices[0].message.content)

cURL

curl https://llm.wavespeed.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
    "model": "moonshotai/kimi-k2.6",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Notes

  • Model: moonshotai/kimi-k2.6
  • Provider: moonshot
  • Best suited for long-horizon coding, UI/UX generation, visual understanding, tool use, multi-agent orchestration, and autonomous workflow execution

Info

Provedormoonshot
Tipollm

Funcionalidades suportadas

Entrada
TextoImagem
Saída
Texto
Contexto262,144
Saída máx.262,142
Vision✓ Suportado
Function Calling✓ Suportado

Guia de acesso à API

Base URLhttps://llm.wavespeed.ai/v1
API Endpointchat/completions
ID do modelomoonshotai/kimi-k2.6

Kimi K2.6 API

moonshotai/kimi-k2.6

Kimi K2.6 is Moonshot AI’s open-source native multimodal agentic model, designed for long-horizon coding, coding-driven UI/UX generation, proactive autonomous execution, and multi-agent orchestration. Built on a 1T-parameter Mixture-of-Experts architecture with 32B active parameters, it supports text and image inputs, a 262K-token context window, thinking mode, preserve-thinking workflows, function calling, and structured outputs. It is especially strong for complex end-to-end coding tasks across Python, Rust, Go, front-end engineering, DevOps, performance optimization, and agentic workflow automation.

Entrada

$0.95 /M

Saída

$4 /M

Contexto

262K

Saída máx.

262K

Vision

Suportado

Uso de ferramentas

Suportado

Experimente Kimi K2.6 no WaveSpeedAI

Acesse Kimi K2.6 através da nossa API unificada — compatível com OpenAI, sem inicializações a frio, preços transparentes.

Perguntas frequentes sobre Kimi K2.6

Quanto custa Kimi K2.6 via API?+

Preços no WaveSpeedAI: $0.95 por milhão de tokens de entrada e $4.00 por milhão de tokens de saída. Prompt caching e batch processing são cobrados separadamente e reduzem o custo efetivo em cargas longas e repetitivas.

Qual é a janela de contexto do Kimi K2.6?+

Kimi K2.6 suporta até 262K tokens de contexto e até 262K tokens de saída por requisição.

Kimi K2.6 é compatível com OpenAI?+

Sim. O WaveSpeedAI expõe o Kimi K2.6 através de um endpoint compatível com OpenAI em https://llm.wavespeed.ai/v1. Aponte o SDK oficial da OpenAI para esta base URL com sua chave API do WaveSpeedAI — sem outras alterações no código.

Como começo a usar o Kimi K2.6?+

Entre no WaveSpeedAI, crie uma chave API em Access Keys, então envie uma requisição para https://llm.wavespeed.ai/v1/chat/completions com o model id mostrado acima. Contas novas recebem créditos grátis para avaliar o Kimi K2.6.

APIs LLM relacionadas