Meet Kimi K3, Moonshot AI's most advanced AI model yet. With 2.8 trillion parameters and a massive one-million-token context window, Kimi K3 can handle complex tasks, large codebases, long documents, and deep research in a single workflow.
Built-in AI agents work directly inside the chat, allowing Kimi to plan, reason, and complete multi-step tasks from start to finish. Whether you are coding, researching, writing, analysing information, or building autonomous agent workflows, Kimi AI provides one powerful platform for everything.
The platform also includes K2.7 Code HighSpeed delivering coding performance at speeds of up to 260 tokens per second - and K2.7 Code is now available in GitHub Copilot.
Seven major releases in thirteen months - each one pushing a specific capability frontier. K3 is the new 2.8T-parameter flagship; the K2 family remains the workhorse lineup.
Kimi K3 is Moonshot AI's next-generation flagship model, built to evolve with the way you work. Powered by 2.8 trillion parameters (Stable LatentMoE - 896 experts, 16 active) and Kimi Delta Attention for 6.3× faster decoding at 1M tokens, it understands long documents, analyses entire codebases, and maintains context across complex conversations.
Native text, image, and video input. Always-on max reasoning. #1 on LMArena Frontend Code Arena (1,679 Elo), 93.5% GPQA Diamond - best of any open-weight model. SWE Marathon 42.0, beating Claude Opus 4.8 and Fable 5. Open weights ship July 27, 2026 under Modified MIT License.
The same Kimi K2.7 Code model - Moonshot's most capable coding model - served at extreme throughput. Up to 260 tokens per second on short-context tasks, ~180 tok/s on median coding inputs. Designed for agentic workflows where speed determines task completion time. Rolling out to Kimi Code Beta, API developers, and Business users.
⚡ 260 tok/s peak · 6× faster $1.90 / $8.00 per 1MMoonshot's most capable coding model in the K2 family. Reduces reasoning token usage by ~30% vs K2.6 while improving scores on every benchmark. Mandatory thinking mode, preserve-thinking across turns. MCP tool-use SOTA: 81.1 on MCP Mark Verified, beating Claude Opus 4.8 at 76.4. Multimodal via MoonViT 400M encoder. Since July 1, 2026: the first open-weight model in GitHub Copilot's model picker (Pro/Pro+/Max, Azure-hosted). Open weights under Modified MIT license.
+21.8% Kimi Code Bench v2 −30% reasoning tokens 🐙 In GitHub CopilotLong-horizon agentic coding, general availability. 300-agent swarm with 4,000 coordinated steps. Claw Groups for cross-model collaboration. Document-to-Skill conversion. 262K context window. SWE-bench Verified 80.2% - SOTA among open-source models. BrowseComp Swarm 86.3%. Supports instant mode and thinking mode - the only Kimi model with switchable thinking. The cheapest frontier model in the lineup at $0.55/$2.65 per 1M.
80.2% SWE-bench Verified 300 agents · 4,000 stepsNative multimodal intelligence trained on 15T mixed visual + text tokens. MoonViT 400M encoder for images and video. 256K context. Agent Swarm v1: 100 parallel sub-agents, 1,500 tool calls, 4.5× execution speedup. SWE-bench 76.8%, AIME 96.1%, VideoMMU 86.6%. ⚠ API retirement notice: new users blocked since July 17, 2026; full sunset August 31, 2026. Migrate to K2.7 Code or K3.
100 sub-agents · 4.5× faster ⚠ API retiring Aug 31, 2026Post-trained reasoning variant with interleaved chain-of-thought and native tool use. 200–300 sequential tool calls without losing task context. Native INT4 quantization via QAT for 2× speed vs FP16 - no accuracy loss. Tencent CodeBuddy integrates K2 Thinking as its core engine. Pioneered the think→act→observe→think loop at production scale. Legacy API slugs reached EOL May 25, 2026 - self-host via open weights.
300 tool calls · 2× speed INT4The original trillion-parameter open-source frontier. 1T MoE, 32B active, 128K context, MuonClip optimizer for zero training instability across 15.5T tokens. Set the open-source agentic baseline: SWE-bench 65.8%, MMLU-Pro 73.3%, τ²-bench 80%. Modified MIT License. Still a strong cost-efficient self-hosted model for text-only workflows.
Open source · Modified MITK3's 1,048,576-token context window fits an entire 800K-line codebase in one session. Kimi Delta Attention (KDA) makes it usable - 6.3× faster decoding at 1M tokens - and the API charges no long-context surcharge: flat $3/$15 per 1M across the full window.
Up to 260 tokens per second on short-context tasks, ~180 tok/s on median coding inputs. 6× faster than standard K2.7 Code. Same model, optimized serving infrastructure. Now rolling out to Beta users.
K2.6 coordinates up to 300 sub-agents executing 4,000 steps simultaneously. K3 adds a dedicated kimi-k3-swarm-max variant for swarm workloads. Compress hours of parallel research and code generation into minutes. Available from Allegretto plan upward.
K3 and K2.7 Code always reason before responding - no shortcutting on complex tasks. Preserve-thinking keeps the chain across multi-turn sessions. K3's reasoning_effort runs at max; K2.7 uses 30% fewer reasoning tokens than K2.6 for lower cost per session.
MoonViT 400M encoder processes images alongside text in K2.5, K2.6, and K2.7 Code. K3 goes further with native video input - upload screen recordings, design walkthroughs, or UI bug reproductions. Design-to-code from any visual input.
K2.7 Code scores 81.1 on MCP Mark Verified - beating Claude Opus 4.8 at 76.4. Since July 1, 2026 it's also the first open-weight model in GitHub Copilot's model picker (Azure-hosted, no Moonshot key needed). Reliable GitHub, Notion, Filesystem, Postgres, and Playwright tool invocation.
Direct AI-native access to World Bank datasets, financial market data, economic indicators, and academic research - queryable in plain language inside Kimi workflows. Available from Adagio (200 requests/mo).
Migrate from GPT or Claude by changing two lines: base_url and model. Full function-calling, streaming, structured outputs, and tool use. K3 note: no temperature/top_p/seed - sampling is fixed server-side. Token-based billing separate from membership.
All K2 models ship as open weights on HuggingFace under Modified MIT License. K3 weights arrive July 27, 2026. Deploy with vLLM, SGLang, KTransformers, or TensorRT-LLM. Self-hosting is the GDPR-compliant path for EU workloads - the hosted API routes via China.
Start free with Kimi K3 access, 200 Professional Data requests/month, Kimi Slides, and Deep Research - no credit card required.
Access all models, tools, and agent features. Free to start.
Start immediately, no credit card needed. You get:
The recommended starting point for regular professional use:
Token-based API billing separate from membership. Get a key at platform.kimi.ai. K3: $3.00/1M input · $15.00/1M output. K2.6: $0.55/1M · $2.65/1M. K2.7 Code: $0.95/1M · $4.00/1M. K2.7 HighSpeed: $1.90/1M · $8.00/1M. 🎁 10–30% bonus credits on top-ups through Aug 11.
Write, convert, review, and translate documents. Upload PDFs, Word files, slides, or spreadsheets - get clean structured outputs with summaries and professional formatting. Powered by K2.6 document agent. LaTeX PDF support.
Try Kimi Docs →Two creation modes: Adaptive (30–60 min, research-first with K2 Thinking) and Visual (5–10 min with Nano Banana Pro / Gemini 3 Pro Image). Chart parsing converts static chart images to native editable PPTX objects. Agentic Slides converts any document to a deck.
Try Kimi Slides →Build formulas, pivots, and dashboards from plain language. Process up to 1M rows of data in a single session. Generate charts, financial models, pivot tables, and data summaries. Export to Excel-compatible formats.
Try Kimi Sheets →Create modern, responsive websites from a prompt. Visual Coding with K3 turns design screenshots directly into production-ready code - K3 is #1 on LMArena Frontend Code Arena. Full-stack generation: frontend, backend, auth, and database ops in one session.
Try Website Builder →Terminal-first coding agent powered by K2.7 Code, with K3 available via /model k3. Works in VS Code, Cursor, JetBrains, and Zed. Autonomous coding sessions, codebase navigation, multi-file editing. Starting at $19/month. K2.7 Code HighSpeed Beta available.
Try Kimi Code →24/7 cloud-based AI agent - no server setup needed. Persistent long-term memory. 5,000+ ClawHub skills. 40GB cloud storage. Scheduled automation and 24/7 task execution. Available on Allegretto plan and above.
Try Kimi Claw →Fully OpenAI and Anthropic SDK compatible. Change two lines to switch from GPT or Claude. Token-based billing, separate from app membership. K3 now live.
kimi-k3 - ✦ Flagship · 1M ctx · NEW
kimi-k3-swarm-max - K3 agent swarm
kimi-k2.7-code-highspeed - 260 tok/s ⚡
kimi-k2.7-code - Coding specialist
kimi-k2.6 - General purpose GA
kimi-k2.5 - ⚠ Retiring Aug 31, 2026
kimi-k2-* legacy slugs - EOL May 25, 2026
Cached input: $0.19/1M (K2.x) · $0.30/1M (K3). No long-context surcharge on K3. 🎁 10–30% bonus credits on top-ups through Aug 11. Get API key →
K3: thinking always on · reasoning_effort="max" only · no temperature/top_p/seed · images base64 or ms:// only
K2.7 Code: thinking always on · temperature=1.0
K2.6: thinking switchable (on/off)
K2.5: thinking switchable · ⚠ retiring
Access chain-of-thought via reasoning_content in response. Pass it back in multi-turn sessions to preserve thinking context. K3 default max_tokens is 131,072 - set an explicit lower limit.
Download from HuggingFace → moonshotai. K3 weights arrive July 27, 2026 (Modified MIT). Deploy with vLLM, SGLang, KTransformers, or TensorRT-LLM. Block-fp8 format. Run the Kimi Vendor Verifier before production traffic. A100-class GPU minimum.
Access Kimi K3 on web and mobile - no setup, no credit card. Use Agent mode for multi-step tasks; K3's thinking is always on. Free tier includes 200 Professional Data requests/month, Kimi Slides Adaptive mode, and 256K context.
Get your key at platform.kimi.ai. Set base_url="https://api.moonshot.ai/v1" and model="kimi-k3" in any OpenAI SDK. Works with LangChain, LlamaIndex, and Anthropic SDKs. Coding default: kimi-k2.7-code.
Terminal-first coding agent at kimi.com/code. Works in VS Code, Cursor, JetBrains, and Zed. Default model: K2.7 Code. Switch to K3 with /model k3. Also in GitHub Copilot's model picker since July 1. Join Beta at kimi.com/code/beta for HighSpeed.
Download from huggingface.co/moonshotai. K3 weights arrive July 27, 2026. Deploy with vLLM, SGLang, KTransformers, or TensorRT-LLM. Block-fp8 format. Modified MIT License - commercial use permitted.
Kimi uses music tempo-inspired plan names. All tiers include Kimi K3 access - full 1M context unlocks at Allegretto. API billing is always separate from membership.
All plans include Kimi K3 access · Full 1M context on Allegretto and above. API billing is always separate from membership. Full pricing details →
Kimi K3 ✦ Moonshot AI · Jul 2026 |
Kimi K2.7 Code Moonshot AI · Jun 2026 |
Kimi K2.6 Moonshot AI · Apr 2026 |
GPT-5.6 Sol OpenAI |
Claude Opus 4.8 Anthropic |
Claude Fable 5 Anthropic |
Gemini 2.5 Pro Google |
|
|---|---|---|---|---|---|---|---|
| Architecture & Access | |||||||
| Open weights | ✓ Jul 27 | ✓ MIT | ✓ MIT | ✗ | ✗ | ✗ | ✗ |
| Parameters (total / active) | 2.8T / ~50B | 1T / 32B | 1T / 32B | Undisclosed | ~200B | Undisclosed | ~1T est. |
| Context window | 1M (1,048,576) | 256K | 262K | 1M | 200K | 1M | 1M |
| Multimodal (image/video) | ✓ + native video | ✓ MoonViT | ✓ | ✓ | ✓ | ✓ | ✓ |
| Benchmarks (mix of vendor + independent) | |||||||
| SWE Marathon | 42.0 ★ | - | - | - | 40.0 | 35.0 | - |
| GPQA Diamond | 93.5% ★ | - | - | ~93% | ~91% | ~94% | ~88% |
| LMArena Frontend Code | #1 · 1,679 Elo ★ | - | - | Top 5 | Top 5 | Top 3 | Top 10 |
| SWE-bench Verified | Pending (indep.) | Pending | 80.2% ✓ | ~60% | 64.3%* | - | ~56% |
| MCP Mark Verified | - | 81.1 ★ | ~73 | - | 76.4 | - | - |
| Speed & Cost | |||||||
| API input $/1M | $3.00 | $0.95 | $0.55 | ~$5.00 | $15.00 | ~$20.00 | $1.25 |
| API output $/1M | $15.00 | $4.00 | $2.65 | ~$60.00 | $75.00 | ~$50.00+ | $5.00 |
| Cost per typical task | ~$0.94 | Lower | Lowest | ~$1.04+ | ~$1.80 | ~$3.20 | ~$0.30 |
| HighSpeed serving | Always thinking | 260 tok/s ⚡ (HS) | ~60 tok/s | ~80 tok/s | ~60 tok/s | ~50 tok/s | ~80 tok/s |
| Platform & Ecosystem | |||||||
| Agent Swarm (parallel agents) | Swarm Max | via K2.6 | 300 agents | Limited | Limited | Limited | Limited |
| In GitHub Copilot picker | - | ✓ First open-weight | - | ✓ | ✓ | ✓ | ✓ |
| Integrated office tools | Full suite | Docs/Slides/Sheets/Code | Full suite | Canvas | Artifacts | Artifacts | Workspace |
| Free tier with agent mode | ✓ 3/day | ✓ | ✓ | Limited | Limited | Limited | Limited |
*SWE-bench Pro figure. ★ = vendor-published; independent verification pending for K3 launch-week numbers. Throughput and competitor pricing approximate. Always verify current benchmarks and pricing on official pages.