Kimi K3
- Provider
- Moonshot AI
- Status
- available
- Context
- 1,048,576 tok
- Price
- $3 / $15 /MTok
Kimi K3 is Moonshot AI’s new flagship, released on 16 July 2026 and the first Chinese model to land inside the frontier pack on independent testing. It scores 57.1 on the Artificial Analysis Intelligence Index v4.1 — fourth of all models, behind only Claude Fable 5 (59.9) and GPT-5.6 Sol (58.9), and ahead of Claude Opus 4.8 (55.7) (Artificial Analysis, the Decoder).
Architecturally it is a 2.8-trillion-parameter sparse mixture-of-experts model — 16 of 896 experts active per token — with a 1-million-token context window, native vision, and two new efficiency mechanisms Moonshot calls Kimi Delta Attention and Attention Residuals, which it credits with up to 6.3x faster decoding (kimi.com). On 26 July 2026 Moonshot shipped the open weights — at 2.8T parameters, the largest open-weight model ever released — under the Kimi K3 License, so K3 is now self-hostable and available on third-party hosts and agents including Devin. The one caveat that remains: at $3 / $15 per million tokens hosted it costs roughly triple its predecessor — the end of super-cheap Chinese AI (the Decoder).
Quick specs
| Provider | Moonshot AI |
| Released | 16 July 2026 |
| Status | Available (hosted + open weights, released 26 July 2026) |
| Architecture | Sparse MoE — 2.8T total, 16 of 896 experts active |
| Context window | 1,048,576 tokens (1M) |
| Modalities | Text, image and video in; text out |
| Reasoning | Thinking always on; max effort only at launch |
| Intelligence Index (independent) | 57.1 — 4th overall, ahead of Claude Opus 4.8 |
| Price | $3 in / $15 out per MTok ($0.30 cached input) |
| Best for | Long-horizon agentic coding and knowledge work at half frontier cost |
| Limitations | Enormous to self-host (~594GB); verbose; raised hallucination rate |
What Kimi K3 is
Kimi K3 is a frontier-scale agentic reasoning model — built for navigating large repositories, tool use, debugging and long-horizon knowledge work rather than quick chat. Thinking is always on: at launch only the max reasoning-effort level is available, with low and high modes promised in later updates (platform.kimi.ai). The model ships natively quantised (MXFP4 weights, MXFP8 activations), and Moonshot reports roughly 2.5x better scaling efficiency than the Kimi K2 generation.
It launched everywhere at once: kimi.com, the mobile apps (iOS, Android, HarmonyOS), the Kimi Work desktop agent, the Kimi Code CLI, the Kimi API and OpenRouter. In the consumer app, Moderato ($19/month) subscribers and above get Kimi K3 with a 256K context; Allegretto ($39/month) and above unlock the full 1M window (Kimi pricing).
Unlike every prior Kimi flagship, K3 launched hosted-only — but Moonshot shipped the full open weights on 26 July 2026, a day ahead of its target. The release is a ~594GB MXFP4 checkpoint on Hugging Face under the Kimi K3 License, and uses the exact same quantisation as the hosted API, so self-hosted quality matches the API (Moonshot’s co-founders confirmed this in a 27 July r/LocalLLaMA AMA). That makes K3 the largest open-weight model released to date, and it is now reachable well beyond Moonshot’s own surfaces: on OpenRouter, Together AI, Modal and Cloudflare Workers AI, and inside agent tools including Devin Desktop and CLI — where Cognition called it the first open-source model to approach frontier performance on its FrontierCode 1.1 eval — and, via an OpenRouter key, Cursor.
Benchmark performance
Moonshot’s own launch table (a ceiling) has Kimi K3 trailing Claude Fable 5 and GPT-5.6 Sol while mostly beating Claude Opus 4.8 and GPT-5.5 — it wins about 7 of the 35 published tests and places second or third in most of the rest (officechai).
| Benchmark | Kimi K3 | Claude Fable 5 | GPT-5.6 Sol | Source |
|---|---|---|---|---|
| GPQA Diamond | 93.5% | 92.6% | 94.1% | Vendor |
| Terminal-Bench 2.1 | 88.3% | 84.6% | 88.8% | Vendor |
| BrowseComp | 91.2% | 88.0% | 90.4% | Vendor |
| DeepSWE | 67.5% | 70.0% | 73.0% | Vendor |
| FrontierSWE | 81.2% | 86.6% | 71.3% | Vendor |
| Intelligence Index v4.1 | 57.1 | 59.9 | 58.9 | Independent (AA) |
| GDPval-AA v2 (Elo) | 1668 | 1760 | 1748 | Independent (AA) |
Independent testing broadly confirms the vendor story. On Artificial Analysis, Kimi K3 debuts at 57.1 — fourth overall, above Claude Opus 4.8 (55.7), Grok 4.5 (53.8) and GLM-5.2 (51.1). On GDPval-AA v2 (economically valuable knowledge work) it jumps to an Elo of 1668 from Kimi K2.6’s 1190, passing Claude Opus 4.8 (1600), GLM-5.2 (1514) and GPT-5.5 (1494). It took the top spot on AutomationBench-AA (53%) and leads Arena.ai’s Frontend Code arena, ahead of Claude Fable 5 (Simon Willison).
Two independent findings cut the other way. Artificial Analysis measured a hallucination rate of 51%, up on its predecessor, even as accuracy improved (the Decoder). And the model is verbose: it generated 130M output tokens across the AA evaluation against a 63M average, so real-world costs depend heavily on reasoning-token burn. No SWE-bench Pro or SWE-bench Verified figures had been published at launch — Moonshot’s coding claims rest on newer suites (DeepSWE, FrontierSWE, SWE Marathon) that are not yet independently replayable. See best AI models for the standings.
Pricing and access
Kimi K3 costs $3.00 per million input tokens, $15.00 output, and $0.30 for cached input on the Kimi API and OpenRouter — Sonnet-class pricing, and roughly triple Kimi K2.6’s $0.95 / $4.00. Moonshot claims cache-hit rates above 90% on coding workloads. Despite the sticker increase, Artificial Analysis puts its cost per task at $0.94 — about half Claude Opus 4.8’s $1.80 and in line with GPT-5.6 Sol ($1.04). In the Kimi app it is on paid tiers from $19/month (256K context) with the full 1M window from $39/month. Since 26 July the open weights (~594GB, MXFP4) can also be self-hosted — with vLLM, or through third-party hosts like Together, Modal and Cloudflare Workers AI — though at 2.8T parameters this is a serious-hardware, multi-GPU undertaking, not a run-it-on-your-laptop model.
How Kimi K3 compares
- vs the closed frontier — On the independent Intelligence Index, Kimi K3 (57.1) sits between GPT-5.6 Sol (58.9) and Claude Opus 4.8 (55.7), with Claude Fable 5 (59.9) still clearly on top. No Chinese model has ranked this high before.
- vs other Chinese and open models — It leads GLM-5.2 (51.1), DeepSeek V4 (44.3) and every other open-weights model on the same index. With the weights now shipped (26 July), it is the strongest open-weight model available, by a wide margin — displacing DeepSeek V4 (MIT) as the most capable model you can download, though DeepSeek remains far cheaper and easier to actually run.
- vs its predecessor — Against Kimi K2.6, K3 is nearly three times the size (2.8T vs 1T), quadruples the context (1M vs 256K), and jumps +478 Elo on GDPval-AA — at three times the API price.
Known limitations
Enormous to self-host. The open weights shipped 26 July 2026, but at 2.8T parameters the MXFP4 checkpoint is ~594GB — a multi-GPU, serious-hardware deployment, not a consumer-GPU model, so most users will still reach it through a host. Raised hallucination rate (51%, Artificial Analysis) despite better accuracy. Verbose reasoning — always-on thinking at max effort burns tokens; Simon Willison’s single pelican-SVG test consumed about 25 cents (simonwillison.net). No standardized coding benchmark (SWE-bench Pro/Verified) at launch. China data residency applies to the hosted app and API. Moonshot itself acknowledges a UX gap versus Claude Fable 5 and GPT-5.6 Sol, sensitivity to preserved thinking history, and “excessive proactiveness” on ambiguous tasks (kimi.com).
FAQ
What is Kimi K3?
Kimi K3 is Moonshot AI’s flagship model, released on 16 July 2026: a 2.8-trillion-parameter sparse mixture-of-experts reasoning model with a 1-million-token context window and native vision. It is the highest-ranked Chinese model ever on independent testing — fourth on the Artificial Analysis Intelligence Index, ahead of Claude Opus 4.8.
Is Kimi K3 open source?
Yes, as of 26 July 2026. Moonshot released the full weights — a ~594GB MXFP4 checkpoint on Hugging Face under the Kimi K3 License (a permissive, modified-MIT-style licence) — making it the largest open-weight model released to date. It can be self-hosted (vLLM) or run through third-party hosts such as OpenRouter, Together, Modal and Cloudflare Workers AI, and it’s now in agent tools like Devin. At 2.8T parameters, though, self-hosting needs serious multi-GPU hardware, so most people will still use a host. Confirm the exact licence terms on the model card before commercial use.
How much does Kimi K3 cost?
$3.00 per million input tokens, $15.00 per million output tokens, and $0.30 per million for cached input via the Kimi API and OpenRouter — roughly triple Kimi K2.6’s pricing. In the Kimi app it requires a paid plan: Moderato ($19/month) for a 256K context, Allegretto ($39/month) and above for the full 1M window.
How good is Kimi K3 at coding?
Strong but not leading. Moonshot’s own figures put it at 88.3% on Terminal-Bench 2.1 (level with GPT-5.6 Sol) and ahead of Claude Opus 4.8 on most of its published coding suites, while trailing Claude Fable 5. It leads Arena.ai’s Frontend Code arena. No SWE-bench Pro or Verified score had been published at launch, so its repository-scale coding standing is not yet independently verified.
Is Kimi K3 better than Claude or GPT?
On the independent Artificial Analysis Intelligence Index, Kimi K3 (57.1) ranks above Claude Opus 4.8 (55.7) and below Claude Fable 5 (59.9) and GPT-5.6 Sol (58.9). It is the first Chinese model inside that frontier pack, at roughly half Opus 4.8’s cost per task — but with a higher measured hallucination rate and no open weights yet.
Last verified 28 July 2026. Kimi K3 figures are Moonshot- and third-party-reported (kimi.com, Artificial Analysis, the Decoder, officechai, Simon Willison); the open weights (MXFP4, ~594GB) shipped 26 July 2026 under the Kimi K3 License. Confirm against Moonshot’s official pages and the model card before relying on specific numbers or licence terms.