Kimi K3 is live - 2.8T params · 1M context · Open weights July 27 · 🎁 10-30% bonus API credits through Aug 11 · Join →
Try Kimi K3 Free →
KIMI K3 LIVE · 2.8T PARAMS · 1M CONTEXT · OPEN WEIGHTS JUL 27

Open Agentic
Intelligence

Meet Kimi K3, Moonshot AI's most advanced AI model yet. With 2.8 trillion parameters and a massive one-million-token context window, Kimi K3 can handle complex tasks, large codebases, long documents, and deep research in a single workflow.
Built-in AI agents work directly inside the chat, allowing Kimi to plan, reason, and complete multi-step tasks from start to finish. Whether you are coding, researching, writing, analysing information, or building autonomous agent workflows, Kimi AI provides one powerful platform for everything. The platform also includes K2.7 Code HighSpeed delivering coding performance at speeds of up to 260 tokens per second - and K2.7 Code is now available in GitHub Copilot.

✦ Kimi K3 - #1 LMArena Frontend Code Arena · July 16, 2026 · Native text + image + video
⚡ K2.7 Code HighSpeed - 260 tok/s · Now in GitHub Copilot · First open-weight model in Copilot
U
Audit this 800K-line monorepo, find the auth vulnerabilities, and refactor the session layer for async/await with full tests.
K3
Loading the full repository into my 1M-token context. I'll map the auth flow, flag the vulnerabilities with severity ratings, restructure the session layer, and generate Jest coverage for every path. Starting now…
K3
2.8TParameters (K3)
1MToken Context (K3)
300Max Agents
260tok/s HighSpeed
93.5%GPQA Diamond (K3)
01

All Models

Seven major releases in thirteen months - each one pushing a specific capability frontier. K3 is the new 2.8T-parameter flagship; the K2 family remains the workhorse lineup.

✦ NEW · JULY 16, 2026 · FLAGSHIP · 1M-TOKEN CONTEXT

Kimi K3: Evolves for you

Kimi K3 is Moonshot AI's next-generation flagship model, built to evolve with the way you work. Powered by 2.8 trillion parameters (Stable LatentMoE - 896 experts, 16 active) and Kimi Delta Attention for 6.3× faster decoding at 1M tokens, it understands long documents, analyses entire codebases, and maintains context across complex conversations.
Native text, image, and video input. Always-on max reasoning. #1 on LMArena Frontend Code Arena (1,679 Elo), 93.5% GPQA Diamond - best of any open-weight model. SWE Marathon 42.0, beating Claude Opus 4.8 and Fable 5. Open weights ship July 27, 2026 under Modified MIT License.

✦ 2.8T params · KDA · LatentMoE 1M context · $3/$15 per 1M Open weights Jul 27
Full guide →
⚡ JUNE 15, 2026 · HIGHSPEED

Kimi K2.7 Code HighSpeed

The same Kimi K2.7 Code model - Moonshot's most capable coding model - served at extreme throughput. Up to 260 tokens per second on short-context tasks, ~180 tok/s on median coding inputs. Designed for agentic workflows where speed determines task completion time. Rolling out to Kimi Code Beta, API developers, and Business users.

⚡ 260 tok/s peak · 6× faster $1.90 / $8.00 per 1M
Full guide →
// CODING SPECIALIST · JUNE 12, 2026 · NOW IN GITHUB COPILOT

Kimi K2.7 Code

Moonshot's most capable coding model in the K2 family. Reduces reasoning token usage by ~30% vs K2.6 while improving scores on every benchmark. Mandatory thinking mode, preserve-thinking across turns. MCP tool-use SOTA: 81.1 on MCP Mark Verified, beating Claude Opus 4.8 at 76.4. Multimodal via MoonViT 400M encoder. Since July 1, 2026: the first open-weight model in GitHub Copilot's model picker (Pro/Pro+/Max, Azure-hosted). Open weights under Modified MIT license.

+21.8% Kimi Code Bench v2 −30% reasoning tokens 🐙 In GitHub Copilot
Learn more →
// AGENT SWARM GA · APRIL 20, 2026

Kimi K2.6

Long-horizon agentic coding, general availability. 300-agent swarm with 4,000 coordinated steps. Claw Groups for cross-model collaboration. Document-to-Skill conversion. 262K context window. SWE-bench Verified 80.2% - SOTA among open-source models. BrowseComp Swarm 86.3%. Supports instant mode and thinking mode - the only Kimi model with switchable thinking. The cheapest frontier model in the lineup at $0.55/$2.65 per 1M.

80.2% SWE-bench Verified 300 agents · 4,000 steps
Learn more →
// VISUAL AGENTIC · JAN 27, 2026

Kimi K2.5

Native multimodal intelligence trained on 15T mixed visual + text tokens. MoonViT 400M encoder for images and video. 256K context. Agent Swarm v1: 100 parallel sub-agents, 1,500 tool calls, 4.5× execution speedup. SWE-bench 76.8%, AIME 96.1%, VideoMMU 86.6%. ⚠ API retirement notice: new users blocked since July 17, 2026; full sunset August 31, 2026. Migrate to K2.7 Code or K3.

100 sub-agents · 4.5× faster ⚠ API retiring Aug 31, 2026
Learn more →
// REASONING · NOV 2025

Kimi K2 Thinking

Post-trained reasoning variant with interleaved chain-of-thought and native tool use. 200–300 sequential tool calls without losing task context. Native INT4 quantization via QAT for 2× speed vs FP16 - no accuracy loss. Tencent CodeBuddy integrates K2 Thinking as its core engine. Pioneered the think→act→observe→think loop at production scale. Legacy API slugs reached EOL May 25, 2026 - self-host via open weights.

300 tool calls · 2× speed INT4
Learn more →
// FLAGSHIP OPEN-SOURCE · JULY 2025

Kimi K2

The original trillion-parameter open-source frontier. 1T MoE, 32B active, 128K context, MuonClip optimizer for zero training instability across 15.5T tokens. Set the open-source agentic baseline: SWE-bench 65.8%, MMLU-Pro 73.3%, τ²-bench 80%. Modified MIT License. Still a strong cost-efficient self-hosted model for text-only workflows.

Open source · Modified MIT
Learn more →
02

Key Features

Kimi K3 - 1M Token Context

K3's 1,048,576-token context window fits an entire 800K-line codebase in one session. Kimi Delta Attention (KDA) makes it usable - 6.3× faster decoding at 1M tokens - and the API charges no long-context surcharge: flat $3/$15 per 1M across the full window.

K2.7 Code HighSpeed - 260 tok/s

Up to 260 tokens per second on short-context tasks, ~180 tok/s on median coding inputs. 6× faster than standard K2.7 Code. Same model, optimized serving infrastructure. Now rolling out to Beta users.

🤖

Agent Swarm - 300 Parallel Agents

K2.6 coordinates up to 300 sub-agents executing 4,000 steps simultaneously. K3 adds a dedicated kimi-k3-swarm-max variant for swarm workloads. Compress hours of parallel research and code generation into minutes. Available from Allegretto plan upward.

🧠

Always-On Thinking (K3 & K2.7)

K3 and K2.7 Code always reason before responding - no shortcutting on complex tasks. Preserve-thinking keeps the chain across multi-turn sessions. K3's reasoning_effort runs at max; K2.7 uses 30% fewer reasoning tokens than K2.6 for lower cost per session.

👁

Native Multimodal (Text + Image + Video)

MoonViT 400M encoder processes images alongside text in K2.5, K2.6, and K2.7 Code. K3 goes further with native video input - upload screen recordings, design walkthroughs, or UI bug reproductions. Design-to-code from any visual input.

🔧

MCP Tool Use SOTA + GitHub Copilot

K2.7 Code scores 81.1 on MCP Mark Verified - beating Claude Opus 4.8 at 76.4. Since July 1, 2026 it's also the first open-weight model in GitHub Copilot's model picker (Azure-hosted, no Moonshot key needed). Reliable GitHub, Notion, Filesystem, Postgres, and Playwright tool invocation.

📊

Professional Data - Now Available

Direct AI-native access to World Bank datasets, financial market data, economic indicators, and academic research - queryable in plain language inside Kimi workflows. Available from Adagio (200 requests/mo).

🔗

OpenAI + Anthropic Compatible API

Migrate from GPT or Claude by changing two lines: base_url and model. Full function-calling, streaming, structured outputs, and tool use. K3 note: no temperature/top_p/seed - sampling is fixed server-side. Token-based billing separate from membership.

🏗️

Open Weights - Self-Host

All K2 models ship as open weights on HuggingFace under Modified MIT License. K3 weights arrive July 27, 2026. Deploy with vLLM, SGLang, KTransformers, or TensorRT-LLM. Self-hosting is the GDPR-compliant path for EU workloads - the hosted API routes via China.

03

Sign In / Create Account

Start free with Kimi K3 access, 200 Professional Data requests/month, Kimi Slides, and Deep Research - no credit card required.

04

Kimi Tools

📄

Kimi Docs

Write, convert, review, and translate documents. Upload PDFs, Word files, slides, or spreadsheets - get clean structured outputs with summaries and professional formatting. Powered by K2.6 document agent. LaTeX PDF support.

Try Kimi Docs →
📊

Kimi Slides

Two creation modes: Adaptive (30–60 min, research-first with K2 Thinking) and Visual (5–10 min with Nano Banana Pro / Gemini 3 Pro Image). Chart parsing converts static chart images to native editable PPTX objects. Agentic Slides converts any document to a deck.

Try Kimi Slides →
📈

Kimi Sheets

Build formulas, pivots, and dashboards from plain language. Process up to 1M rows of data in a single session. Generate charts, financial models, pivot tables, and data summaries. Export to Excel-compatible formats.

Try Kimi Sheets →
🌐

Kimi Websites

Create modern, responsive websites from a prompt. Visual Coding with K3 turns design screenshots directly into production-ready code - K3 is #1 on LMArena Frontend Code Arena. Full-stack generation: frontend, backend, auth, and database ops in one session.

Try Website Builder →
💻

Kimi Code CLI

Terminal-first coding agent powered by K2.7 Code, with K3 available via /model k3. Works in VS Code, Cursor, JetBrains, and Zed. Autonomous coding sessions, codebase navigation, multi-file editing. Starting at $19/month. K2.7 Code HighSpeed Beta available.

Try Kimi Code →
🦞

Kimi Claw (OpenClaw)

24/7 cloud-based AI agent - no server setup needed. Persistent long-term memory. 5,000+ ClawHub skills. 40GB cloud storage. Scheduled automation and 24/7 task execution. Available on Allegretto plan and above.

Try Kimi Claw →
05

API Access

Fully OpenAI and Anthropic SDK compatible. Change two lines to switch from GPT or Claude. Token-based billing, separate from app membership. K3 now live.

Model IDs

kimi-k3 - ✦ Flagship · 1M ctx · NEW
kimi-k3-swarm-max - K3 agent swarm
kimi-k2.7-code-highspeed - 260 tok/s ⚡
kimi-k2.7-code - Coding specialist
kimi-k2.6 - General purpose GA
kimi-k2.5 - ⚠ Retiring Aug 31, 2026
kimi-k2-* legacy slugs - EOL May 25, 2026

API Pricing (per 1M tokens)

✦ K3 $3.00 in · $15.00 out
K2.7 $0.95 in · $4.00 out
⚡HS $1.90 in · $8.00 out
K2.6 $0.55 in · $2.65 out

Cached input: $0.19/1M (K2.x) · $0.30/1M (K3). No long-context surcharge on K3. 🎁 10–30% bonus credits on top-ups through Aug 11. Get API key →

Thinking Mode Defaults

K3: thinking always on · reasoning_effort="max" only · no temperature/top_p/seed · images base64 or ms:// only
K2.7 Code: thinking always on · temperature=1.0
K2.6: thinking switchable (on/off)
K2.5: thinking switchable · ⚠ retiring

Access chain-of-thought via reasoning_content in response. Pass it back in multi-turn sessions to preserve thinking context. K3 default max_tokens is 131,072 - set an explicit lower limit.

Self-Host with Open Weights

Download from HuggingFace → moonshotai. K3 weights arrive July 27, 2026 (Modified MIT). Deploy with vLLM, SGLang, KTransformers, or TensorRT-LLM. Block-fp8 format. Run the Kimi Vendor Verifier before production traffic. A100-class GPU minimum.

Python · OpenAI SDK · All Models
# pip install openai from openai import OpenAI client = OpenAI( api_key="YOUR_KIMI_API_KEY", base_url="https://api.moonshot.ai/v1" ) # ── ✦ Kimi K3 (flagship · 1M context) ── response = client.chat.completions.create( model="kimi-k3", max_tokens=8192, # default 131,072 - set explicit messages=[ {"role": "user", "content": "Audit this 800K-line repo..."} ] ) print(response.choices[0].message.reasoning_content) print(response.choices[0].message.content) # ── K2.7 Code HighSpeed (6× faster) ── response = client.chat.completions.create( model="kimi-k2.7-code-highspeed", max_tokens=32768, messages=[ {"role": "system", "content": "You are a senior engineer."}, {"role": "user", "content": "Refactor this auth module..."} ] ) # ── K2.6 with switchable thinking ── response = client.chat.completions.create( model="kimi-k2.6", temperature=0.6, extra_body={"thinking": False}, # instant mode messages=[{"role": "user", "content": "Summarize..."}] ) # ── Preserve thinking for multi-turn agents ── messages = [] for step in agent_steps: messages.append({"role": "user", "content": step}) r = client.chat.completions.create( model="kimi-k3", messages=messages, max_tokens=16384 ) msg = r.choices[0].message messages.append({ "role": "assistant", "content": msg.content, "reasoning_content": msg.reasoning_content # ← preserve })
06

How to Use Kimi AI

01

Start Free at kimi.com

Access Kimi K3 on web and mobile - no setup, no credit card. Use Agent mode for multi-step tasks; K3's thinking is always on. Free tier includes 200 Professional Data requests/month, Kimi Slides Adaptive mode, and 256K context.

02

Integrate via API

Get your key at platform.kimi.ai. Set base_url="https://api.moonshot.ai/v1" and model="kimi-k3" in any OpenAI SDK. Works with LangChain, LlamaIndex, and Anthropic SDKs. Coding default: kimi-k2.7-code.

03

Install Kimi Code CLI

Terminal-first coding agent at kimi.com/code. Works in VS Code, Cursor, JetBrains, and Zed. Default model: K2.7 Code. Switch to K3 with /model k3. Also in GitHub Copilot's model picker since July 1. Join Beta at kimi.com/code/beta for HighSpeed.

04

Self-Host Open Weights

Download from huggingface.co/moonshotai. K3 weights arrive July 27, 2026. Deploy with vLLM, SGLang, KTransformers, or TensorRT-LLM. Block-fp8 format. Modified MIT License - commercial use permitted.

07

Moonshot AI Research

Kimi K3
2.8T flagship, 1M ctx, always-on thinking
KDA
Kimi Delta Attention - 6.3× faster at 1M
Stable LatentMoE
896 experts, 16 active - 2.5× scaling
Kimi K2.7 Code
Coding-specialist, −30% reasoning tokens
Kimi K2.6
General GA, 300-agent swarm, 262K ctx
Kimi K2.5
Visual agentic intelligence, 100 agents
Kimi K2 Thinking
Interleaved reasoning + tool use
Kimi K2
Flagship open-source MoE, 128K ctx
MuonClip
Stable 1T-param MoE training optimizer
MoonViT
400M vision encoder for multimodal
Mooncake
Efficient LLM serving - 90%+ cache hits
Kimi-VL
Multimodal image-text research model
Kimi-Audio
Speech and audio capabilities
Kimina-Prover
Formal logic validation and reasoning
MoBA
Block attention for long-context models
WorldVQA
Vision-centric world knowledge benchmark
08

Membership Pricing

Kimi uses music tempo-inspired plan names. All tiers include Kimi K3 access - full 1M context unlocks at Allegretto. API billing is always separate from membership.

Adagio
Free Forever
$0/mo
Always free. No credit card.
  • Includes
  • Kimi K3 web + mobile (256K ctx)
  • Kimi K2.6 chat access
  • Kimi Slides (Adaptive)
  • Deep Research (limited)
  • 200 Pro Data requests
  • Agent mode (3/day)
Moderato
Best for most users
$19/mo
$180/yr billed annually (save $48)
  • Adagio +
  • Kimi K3 priority queue (256K ctx)
  • Kimi Slides Visual (Nano Banana Pro)
  • 60 agent credits/month
  • 2,000 Pro Data requests
  • Deep Research full access
  • Kimi Code 1× credits
Allegretto
Power users
$39/mo
$372/yr billed annually (save $96)
  • Moderato +
  • Kimi K3 full 1M context ✦
  • 150 agent credits/month
  • Kimi Code 5× credits
  • Kimi Claw cloud access
  • Agent Swarm 50 uses/4 agents
  • 5,000 Pro Data requests
Allegro
Teams & power users
$99/mo
$948/yr billed annually (save $240)
  • Allegretto +
  • Kimi K3 full 1M context ✦
  • 360 agent credits/month
  • Kimi Code 15× credits
  • Agent Swarm 120 uses/6 agents
  • 12,000 Pro Data requests
  • Priority Claw scheduling
Vivace
Enterprise
$199/mo
$1,908/yr billed annually (save $480)
  • Allegro +
  • K3 1M ctx + dedicated queue ✦
  • 720 agent credits/month
  • Kimi Code 30× credits
  • Swarm 240 uses · 8 agents
  • 24,000 Pro Data requests
  • Enterprise SLA available

All plans include Kimi K3 access · Full 1M context on Allegretto and above. API billing is always separate from membership. Full pricing details →

09

How Kimi Compares

Kimi K3 ✦
Moonshot AI · Jul 2026
Kimi K2.7 Code
Moonshot AI · Jun 2026
Kimi K2.6
Moonshot AI · Apr 2026
GPT-5.6 Sol
OpenAI
Claude Opus 4.8
Anthropic
Claude Fable 5
Anthropic
Gemini 2.5 Pro
Google
Architecture & Access
Open weights Jul 27 MIT MIT
Parameters (total / active) 2.8T / ~50B 1T / 32B 1T / 32B Undisclosed ~200B Undisclosed ~1T est.
Context window 1M (1,048,576) 256K 262K 1M 200K 1M 1M
Multimodal (image/video) + native video MoonViT
Benchmarks (mix of vendor + independent)
SWE Marathon 42.0 ★ - - - 40.0 35.0 -
GPQA Diamond 93.5% ★ - - ~93% ~91% ~94% ~88%
LMArena Frontend Code #1 · 1,679 Elo ★ - - Top 5 Top 5 Top 3 Top 10
SWE-bench Verified Pending (indep.) Pending 80.2% ✓ ~60% 64.3%* - ~56%
MCP Mark Verified - 81.1 ★ ~73 - 76.4 - -
Speed & Cost
API input $/1M $3.00 $0.95 $0.55 ~$5.00 $15.00 ~$20.00 $1.25
API output $/1M $15.00 $4.00 $2.65 ~$60.00 $75.00 ~$50.00+ $5.00
Cost per typical task ~$0.94 Lower Lowest ~$1.04+ ~$1.80 ~$3.20 ~$0.30
HighSpeed serving Always thinking 260 tok/s ⚡ (HS) ~60 tok/s ~80 tok/s ~60 tok/s ~50 tok/s ~80 tok/s
Platform & Ecosystem
Agent Swarm (parallel agents) Swarm Max via K2.6 300 agents Limited Limited Limited Limited
In GitHub Copilot picker - First open-weight -
Integrated office tools Full suite Docs/Slides/Sheets/Code Full suite Canvas Artifacts Artifacts Workspace
Free tier with agent mode 3/day Limited Limited Limited Limited

*SWE-bench Pro figure. ★ = vendor-published; independent verification pending for K3 launch-week numbers. Throughput and competitor pricing approximate. Always verify current benchmarks and pricing on official pages.

10

FAQ

Open Intelligence,
Next-Level Performance

Kimi K3 is now live with 2.8 trillion parameters, a one-million-token context window, and powerful agent capabilities built directly into chat. Handle complex coding, research, writing, and multi-step tasks in one seamless workflow.
Open weights ship July 27. K2.7 Code is now in GitHub Copilot. 🎁 10–30% bonus API credits through August 11. Free to try.