Moonshot AI Releases Kimi K3: 2.8T Parameter Open-Weight Multimodal Frontier Model

Moonshot AI has released Kimi K3, a 2.8-trillion parameter MoE powerhouse featuring 1M context, native video understanding, and leading mathematical reasoning.
| Platform | Route Identifier | Context | Max Output | Rate / Tier | Status |
|---|---|---|---|---|---|
| Moonshot Open Platform | kimi-k3 | 1.05M tokens (1,048,576) | 944K tokens (943,718) | $3.00 / $15.00 per 1M tokens ($0.30 cached prompt) | active |
| OpenRouter | moonshotai/kimi-k3 | 1.05M tokens (1,048,576) | 944K tokens (943,718) | $2.30 / $11.55 per 1M tokens ($0.26 cached prompt) | active |
Key Takeaways
- check_circleKimi K3 scales Mixture-of-Experts to 2.8 trillion total parameters, achieving near-frontier reasoning on complex scientific and coding benchmarks.
- check_circleNative video and image perception enables high-accuracy frame-by-frame scene analysis across 1M context tokens.
- check_circlePriced on the Moonshot Open Platform at $3.00/M prompt and $15.00/M completion, with the OpenRouter route at $2.30/M and $11.55/M.
- check_circleAIMI detected the live production route on 16 July 2026 at 17:30:58 SAST.
Massive Scale Mixture-of-Experts
Moonshot AI has cemented its position at the forefront of global AI research with Kimi K3. Utilizing an unprecedented 2.8T parameter MoE architecture with top-k expert activation, Kimi K3 achieves frontier general intelligence while maintaining inference efficiency comparable to much smaller dense models.
The model showcases remarkable capabilities in complex mathematical theorem proving, financial modeling, and Chinese-English bilingual translation.
Frequently Asked Questions
How large is Moonshot AI Kimi K3?
Kimi K3 is a 2.8-trillion parameter Mixture-of-Experts (MoE) model designed for 1M context lengths.
Related Frontier Models & Releases
Explore verified specifications, benchmark results, and route pricing across alternative models in this class.
DeepSeek releases V4.1 Flash with native multimodal vision, 552B MoE and lower API rates
DeepSeek has launched DeepSeek-V4.1-Flash, a 552B-parameter mixture-of-experts model featuring a novel Causal Encoder–Decoder architecture with just 8B active input and 16B active output parameters. The release brings native vision, compresses KV cache storage by up to 8x, and slashes off-peak API pricing to $0.15 per million input tokens.
Alibaba Qwen Releases Qwen3.8 Max: 2.4-Trillion Parameter Flagship with 1M Multimodal Context
Alibaba's Qwen team has launched Qwen3.8 Max, a 2.4T parameter Mixture-of-Experts model offering 1M token context, native video perception, and deep agent tool orchestration.
Meta Releases Muse Spark 1.3: 1M Multimodal Reasoning Model for Autonomous Agent Workflows
Meta Superintelligence Labs has published Muse Spark 1.3, an open-weights frontier agent model with 1,048,576 context tokens and comprehensive multimodal comprehension.