Five plans from free to $199/month. K3 is now the flagship - accessible in the app across all tiers, with full 1M context unlocking at Allegretto. API billing always separate from membership.
K3 is live. Get 10–30% bonus credits on all prepaid API top-ups at platform.kimi.ai through August 11, 2026. One-time top-up of ¥199 activates API access.
All plans support 256K context minimum (1M ctx on Allegretto+). API billing is always separate from membership. Prices in USD. Taxes not included. Quotas subject to change - verify at platform.kimi.ai.
| Feature | AdagioFree | Moderato$19/mo | Allegretto$39/mo | Allegro$99/mo | Vivace$199/mo |
|---|
API billing is completely separate from your membership plan - token-based, pay-as-you-go. One API key unlocks all available models. Get your key at platform.kimi.ai.
https://api.moonshot.ai/v1base_url and model from any existing clientopenrouter.ai/moonshotai/kimi-k3 at same ratesMoonshot has issued a retirement notice for kimi-k2.5 and moonshot-v1 on August 31, 2026. Migration to kimi-k2.7-code (routine coding) or kimi-k3 (max quality) is recommended. K2.7 Code is not in the retirement notice.
All tiers get access to Kimi K3 in the app. What changes by plan is the context window available, the agent credit quota, and Kimi Code multiplier.
kimi-k2.7-code-highspeedShort answer: free to start, $19–$199/month for membership, and token-based API from $0.30/1M (cached) to $15/1M. Here is the full picture in one place.
| Use case | Best option | Typical monthly cost |
|---|---|---|
| Casual chat and research | Adagio (Free) | $0 |
| Daily professional use | Moderato $19 + API | $25–$60 |
| Developer building with API | API only (K2.6 or K2.7) | $5–$200 (usage-based) |
| Agent swarm workflows | Allegretto $39 + API | $60–$200 |
| K3 - frontier quality tasks | API kimi-k3 | ~$0.94/task |
| High-volume AI infrastructure | Vivace $199 + API + self-host | $500–$5,000+ |
A single reference table covering every plan dimension - monthly and annual rates, context access, model credits, tools, and quotas.
| Dimension | AdagioFree | Moderato$19/mo | Allegretto$39/mo | Allegro$99/mo | Vivace$199/mo |
|---|
Released July 16, 2026. The world's largest open-weight model at 2.8T parameters - priced at the same tier as Claude Sonnet, but delivering frontier-level coding performance.
K3 has no non-thinking mode. Every response includes a reasoning trace billed at the full $15.00/1M output rate. A chatty thinking trace (common on complex tasks) can cost more than the visible answer. K3 is also the only SKU - there is no cheaper "instant" K3 variant.
Released June 12–15, 2026. The default model for Kimi Code CLI. Two speed tiers: standard and HighSpeed (Beta, 5–6× faster model output at 2× the token price).
General availability since April 20, 2026. The cheapest frontier-class model in the Kimi lineup - and the one to use when cost matters for agent swarm and multi-step workflows.
Tokens are the unit of all Kimi API billing. Everything - input, output, reasoning traces, system prompts, conversation history, tool calls, and image descriptions - is measured in tokens.
System prompt + conversation history + your message + any uploaded images/video descriptions. Billed once per request. Cache hits dramatically reduce this cost.
Visible answer + thinking trace (on K2.7 Code and K3). Always-on thinking models produce thinking tokens at the full output rate - this is where K3 costs can surprise you.
When Mooncake serving recognizes repeated content (your system prompt, repo context, document), the input cost drops to $0.19/1M (K2.x) or $0.30/1M (K3). Cache hit rates on coding tasks exceed 90%.
Tool invocations (file reads, shell commands, web scraping) consume tokens when results are returned to the model. Web search costs an additional flat $0.015/call.
These are separate systems that serve different use cases. Most power users need both. Here is how to decide.
Token prices are simple - what makes bills unpredictable are the multipliers. Here are the seven biggest cost drivers across Kimi models.
K3 and K2.7 Code always produce a reasoning trace before answering. Thinking tokens bill as output at the full output rate ($15/1M for K3, $4/1M for K2.7). On complex tasks the thinking trace can be 2–5× longer than the visible answer.
Every turn sends the full conversation history as input. A 10-turn session where each exchange is 5K tokens means turn 10 pays for 50K+ input tokens. Cache hits reduce this but don't eliminate it - new user messages always count as fresh input.
Each web search invocation costs $0.015 flat, plus the tokens consumed reading results. In a Deep Research session running 100+ searches, that's $1.50 in search calls alone before counting input/output tokens.
Each agent step is a separate API call - with the full accumulated context as input. A 300-step agent session at 10K average context per step = 3M input tokens before any output. Caching helps only for repeated context.
Images are tokenized before processing. A high-res screenshot can consume 2K–4K tokens of input. In a multimodal coding session with 10 UI screenshots, that's 20K–40K tokens of image-only input per request.
If your system prompt or document context changes every request, you pay full fresh-input rates every time. Pin your system prompt and load documents once; let Mooncake's cache handle the rest. This alone can cut input costs by 80–90%.
K3 output costs 5.5× more per token than K2.6 ($15 vs $2.65). For tasks like summarization, classification, or quick code edits where K2.6 or K2.7 produces equivalent results, using K3 is pure cost inflation with no quality benefit.
Skip the comparison tables. Find your profile and get a direct recommendation.
How Kimi's plans and API rates compare to ChatGPT, Claude, Gemini, Perplexity, and Copilot - across subscription price, API cost, open weights, context, and agent capability.
| Dimension | Kimi AIMoonshot AI | ChatGPTOpenAI | ClaudeAnthropic | GeminiGoogle | PerplexityPerplexity AI |
|---|
There is no separate "agent" pricing tier. Agent workflows are billed at the same per-token rates - but agents multiply token consumption in ways that matter for cost planning.
A 300-step Agent Swarm session makes 300+ API calls. Each call sends the full accumulated task context as input. No special "agent" discount - regular token rates apply.
File reads, shell output, web scraping results - all returned to the model as additional input tokens. A file-heavy coding agent reading 50 files across a session adds tens of thousands of input tokens.
Every web search invocation by the agent costs $0.015 flat plus tokens for reading search results. Plan for $1–$3 in search fees per Deep Research session.
Agent Swarm capacity (50–240 swarm uses, 4–8 agents) is included in Allegretto+ membership plans. Using Agent Swarm via the Kimi app does not charge additional API tokens from your API account - it runs against membership credits. Only direct API calls to kimi-k2.6 for agent orchestration are token-billed.
Images and video are billed in tokens - no separate multimodal surcharge. The cost depends on the model and media resolution, not a per-image flat fee.
Images are tokenized based on resolution. A typical UI screenshot at 1280×800 consumes approximately 1,500–3,000 input tokens. Available on K2.5 (via MoonViT), K2.6, K2.7 Code, and K3.
K3 is the only Kimi model supporting native video input. Frames are extracted and each is tokenized as an image. A 60-second clip at 1 frame/second = ~60 image-tokens. Resolution determines per-frame token count.
K3's 1M token context is the largest in the Kimi lineup. Critically, Moonshot charges flat rates with no long-context surcharge - unlike some competitors that charge premium rates above 200K tokens.
Standard transformer attention scales quadratically with context length (O(n²)). At 1M tokens, this is computationally prohibitive. Kimi Delta Attention (KDA) replaces most attention with a linear recurrence, inserting sparse "delta" full-attention passes only where needed.
Web search is billed as a flat tool-call fee per invocation, on top of the regular token charges for reading and processing search results.
web_search tool callweb_search as a tool in the tools arrayweb_search in tools when not needed