Open Frontier Intelligence: Moonshot's Kimi K3 Redefines the Open-Weight Ceiling
Moonshot AI's Kimi K3 ships as a 2.8T-parameter open MoE with 1M context, native vision, and 2.5× scaling efficiency over K2 — the first open 3T-class frontier model.
Moonshot AI's Kimi K3 is not another incremental open-weight release. At 2.8 trillion parameters — with 104 billion activated per token via Stable LatentMoE — it is the first open 3T-class model, pairing a 1-million-token context window with native vision and agentic reinforcement learning. The technical report landed on arXiv July 27; full weights followed on GitHub the same week.
Architecture: KDA, AttnRes, and 16-of-896 routing
The arXiv paper (2607.24653) centers three structural bets:
- Kimi Delta Attention (KDA) — hybrid linear attention that improves information flow across long sequences.
- Attention Residuals (AttnRes) — residual pathways across model depth.
- Stable LatentMoE — routes each token through 16 of 896 experts, yielding roughly 2.5× scaling efficiency over Kimi K2.
Moonshot reports 104B activated parameters per forward pass despite the 2.8T total — a sparse design meant to keep inference tractable while preserving frontier capability.
What the benchmarks claim
Moonshot's GitHub README and tech blog cite frontier-level performance across long-horizon coding, agentic tasks, knowledge, reasoning, and vision. The paper is explicit about the ceiling: Kimi K3 still trails the strongest proprietary models in Moonshot's suite — Claude Fable 5 and GPT-5.6 Sol — but outperforms other open and proprietary models evaluated.
Notable evaluation choices:
- SWE-Marathon and FrontierSWE for software engineering, using Moonshot's Kimi Code harness.
- PostTrainBench at maximum reasoning effort.
- CritPt and AA-LCR scores cited from Artificial Analysis (July 23, 2026).
Infrastructure as a research result
At 2.8T scale, training infrastructure is part of the scientific contribution. Moonshot highlights:
- Algorithm-system co-design for KDA.
- Perfectly balanced expert-parallel training with efficient memory management.
- Million-token agentic RL with persistent rollout and sandbox states.
Post-training spans general, agentic, and coding domains at multiple reasoning-effort levels — a bet that compositional generalization and long-horizon execution require RL at scale, not just pretraining.
Open weights and licensing
The full Kimi K3 weights are released under the Kimi K3 License via MoonshotAI/Kimi-K3. API access is live on Kimi.com, Kimi Work, Kimi Code, and the Kimi API — with max thinking effort as the default at launch.
For the open-model ecosystem, Kimi K3 sets a new parameter watermark. For frontier labs, it raises the bar on what "open" means when trillion-parameter multimodal agents ship with million-token RL infrastructure.
### Sources
- arXiv — Kimi K3: Open Frontier Intelligence (2607.24653) (July 27, 2026)
- GitHub — MoonshotAI/Kimi-K3 (July 27, 2026)
- Moonshot AI — Kimi K3 Tech Blog (July 2026)