Kimi K3: Moonshot's 2.8T frontier model
Moonshot's 2.8-trillion-parameter Mixture-of-Experts model, with a 1M-token context window, native image and video input, and open weights.
1M tokens · Text / Vision / Video / Code · Prompt cache
Okou does not run Kimi K3. This page is a reference for its specs, pricing and benchmarks. For the same kind of work on Okou today, use Claude Fable 5.
See Claude Fable 5Kimi K3 is Moonshot AI's frontier model, launched July 16, 2026: a sparse Mixture-of-Experts network with 2.8 trillion total parameters, 16 of 896 experts active per token, a 1,048,576-token context window and native image and video input. Artificial Analysis scored it 57.1 on the Intelligence Index v4.1 — fourth among tested configurations, behind Claude Fable 5 (59.9) and GPT 5.6 Sol (58.9), ahead of Claude Opus 4.8.
List price is $3.00 per million input tokens, $0.30 on a cache hit and $15.00 per million output tokens, flat across the whole window with no long-context surcharge. Open weights went up on Hugging Face on July 27, 2026 under Moonshot's own Kimi K3 License. Okou does not run Kimi K3. The closest model in the Okou lineup is Claude Fable 5, which matches the 1M-token window and leads the same Intelligence Index.
What is Kimi K3?
July 16, 2026 · Moonshot's frontier line. Successor to the Kimi K2 series, including K2.7 Code.
Kimi K3 is Moonshot AI's frontier model, unveiled on July 16, 2026 and served through the Kimi app and the Kimi API as kimi-k3. It is a sparse Mixture-of-Experts model: 2.8 trillion total parameters with 16 of 896 experts active per token, which is what keeps inference priced near a mid-tier dense model despite the parameter count.
The jump over the K2 line is largest on long-horizon agent work. On Moonshot's launch suite K3 leads every model tested on Program Bench, SWE Marathon, BrowseComp, SpreadsheetBench 2 and Automation Bench, while trailing Claude Fable 5 on FrontierSWE and GPT 5.6 Sol on DeepSWE. On Artificial Analysis's GDPval-AA v2 Elo, which scores economically valuable knowledge work, it reaches 1668, up from Kimi K2.6's 1190 and above Claude Opus 4.8's 1600.
What's notable about Kimi K3
Headline architecture and capability features.
K3 pairs the MoE stack with two efficiency mechanisms Moonshot calls Kimi Delta Attention and Attention Residuals, credited with up to 6.3x faster decoding, and serves its 1,048,576-token window at a single flat price with no context tiering. Output runs to 131,072 tokens by default and can be configured higher within the same window. Input covers text, images and video; output is text. The API at platform.kimi.ai is OpenAI-compatible, and the weights — roughly 594 GB in MXFP4 — are published for self-hosting.
Specs at a glance
Kimi K3 benchmarks
Coding, agent and visual scores are Moonshot's launch numbers, with every model at maximum thinking effort; the Intelligence Index and the GDPval Elo are Artificial Analysis's independent runs. A vendor table is a ceiling rather than a guarantee, and independent testers have reported a hallucination rate well above Moonshot's own figure.
Kimi K3 pricing
Provider list price, per 1M tokens.
How Kimi K3 behaves in practice
Reported behaviour from the vendor's launch material and independent evaluations.
Long-horizon agent runs
The scores K3 leads outright — SWE Marathon, Automation Bench, BrowseComp — are the ones that measure staying on task across hundreds of steps, rather than single-shot correctness. That matches Moonshot's framing of the model as an agent engine first.
Flat 1M-token pricing
The full 1,048,576-token window bills at one rate, with no surcharge past a threshold. GPT-5.5's 1M window, by contrast, prices input above 272K tokens differently, which is what makes very long single-context runs expensive elsewhere.
Cost per task
Artificial Analysis measured $0.94 per task across its index — roughly half Claude Opus 4.8's $1.80 and in line with GPT 5.6 Sol's $1.04. Moonshot cites cache-hit rates above 90% on coding workloads, which pulls effective input cost toward the $0.30 cached rate.
Best agent tasks for Kimi K3
Repository-scale reading in one context
A 1M-token window at a flat price fits an entire codebase, a full quarter of support tickets or a long document set into a single request without chunking, and without a long-context surcharge waiting at the far end of the window.
Browsing and research agents
BrowseComp 91.2 is the highest score in Moonshot's launch table, ahead of GPT 5.6 Sol and Claude Fable 5. Multi-hop web research — find the source, follow the citation, reconcile the numbers — is where K3's agent training shows most clearly.
Screenshot and video input
K3 takes images and video natively, so QA agents reading screenshots and pipelines processing recorded clips do not need a separate vision model spliced in front of the reasoning model.
When to skip Kimi K3
Skip K3 when the work needs a vendor-audited safety profile or contractual enterprise support in the US or EU; when prompts are short enough that the $0.30 cache-hit rate never lands and a cost-saving model would do; and when self-hosting is the point — 594 GB of MXFP4 weights wants 64 or more accelerators for a production deployment. Independent testing has also measured hallucination rates well above Moonshot's published figure, so fact-heavy output still needs checking. And on Okou specifically, K3 is not an option at all.
Kimi K3 vs other models
Kimi K3 vs Claude Fable 5
Fable 5 leads the Intelligence Index (59.9 to 57.1) and FrontierSWE (86.6 to 81.2); K3 takes Terminal-Bench 2.1, BrowseComp and SWE Marathon. Both carry a 1M-token window. What usually decides it is price and availability: Fable 5 is $10/$50 per million tokens and runs on Okou, K3 is $3/$15 and does not.
Kimi K3 vs GPT 5.6 Sol
Sol scores 58.9 on the Intelligence Index to K3's 57.1, and wins DeepSWE (73.0 to 67.5) and Terminal-Bench by half a point. K3 answers with a 1M window against Sol's 400K and a third of the input price. Measured cost per task lands close: $0.94 for K3, $1.04 for Sol.
Kimi K3 vs Claude Opus 4.8
K3 is ahead of Opus 4.8 on most agentic suites and on GDPval-AA Elo (1668 to 1600), at $3/$15 against $5/$25. Opus 4.8 keeps the edge on tool-routing reliability in production and is available on Okou, where K3 is not.
Kimi K3 vs Kimi K2.7 Code
K3 is the successor: 1M tokens of context against 256K, native video input, and a much higher ceiling on agentic benchmarks — at roughly triple the price ($3/$15 against $1.14/$4.80). Neither runs on Okou today; K2.7 Code was delisted, and K3 was never added.
Bottom line: should you use Kimi K3?
The strongest open-weight model released so far, and the first that competes with the closed frontier on agent work at a third of Fable 5's price. Okou does not run it, so read this page as reference: for the same 1M-token agent workloads on Okou, Claude Fable 5 is the closest match.
Frequently asked questions
Can I use Kimi K3 on Okou?
Not today. Kimi K3 is not part of the Okou model lineup — it cannot be picked in chat or in a workflow, and it cannot be added with your own API key. Claude Fable 5 is the closest model Okou runs: the same 1M-token context window, and the top score on the same Intelligence Index.
When was Kimi K3 released?
Moonshot AI launched Kimi K3 on July 16, 2026 in the Kimi app and the Kimi API, and published the open weights on Hugging Face on July 27, 2026.
What is Kimi K3's context window?
1,048,576 tokens, covering input and output together, priced flat across the whole window. Output defaults to 131,072 tokens and can be configured higher within the remaining context.
How much does Kimi K3 cost?
$3.00 per million input tokens on a cache miss, $0.30 per million on a cache hit and $15.00 per million output tokens, with no long-context surcharge. In the Kimi app, paid tiers start at $19/month for a 256K window, and the full 1M window unlocks at $39/month.
Is Kimi K3 open source?
The weights are open under Moonshot's own Kimi K3 License, which reads like MIT with two commercial conditions: a model-as-a-service business needs a separate agreement with Moonshot once revenue passes $20M over any 12 months, and any product above 100 million monthly active users or $20M in monthly revenue has to display "Kimi K3" in its interface.
Alternatives
Availability of Kimi K3 on Okou
Kimi K3 is not part of the Okou model lineup, so it cannot be picked in chat or in a workflow, and it cannot be added with your own API key. Claude Fable 5 is the closest model Okou runs today.