Open Weights at 2.8 Trillion Parameters: Kimi K3 Sets a New Bar for Deployable Frontier Models
Moonshot AI's Kimi K3 tech report on arXiv describes a 2.8T-parameter MoE model with open weights, 1M-token context, and native vision — trailing only Claude Fable 5 and GPT-5.6 Sol on the authors' eval suite.
The open-model race no longer stops at "good enough for fine-tuning." On July 27, 2026, researchers from Moonshot AI posted Kimi K3 to arXiv — a 2.8 trillion-parameter Mixture-of-Experts model with full weight release, native vision, and a one-million-token context window. The authors claim frontier-level performance on long-horizon coding, agentic tasks, and multimodal reasoning while still trailing the strongest closed models.
Architecture in brief
Kimi K3 activates 104 billion parameters per forward pass from a 2.8T total, routing through Stable LatentMoE with 16 of 896 experts active per token. The backbone combines Kimi Delta Attention and Attention Residuals to improve information flow across both sequence length and depth.
The authors report roughly 2.5× better scaling efficiency versus Kimi K2 — a claim that matters because MoE models live or die on routing stability and training economics, not parameter counts alone.
Post-training emphasizes reinforcement learning across general, agentic, and coding domains at multiple reasoning-effort levels. The paper highlights compositional generalization and long-horizon execution — capabilities that show up in agent benchmarks more than static Q&A leaderboards.
Where it lands on evals
Moonshot's evaluation suite places Kimi K3 behind Claude Fable 5 and GPT-5.6 Sol on overall performance, but ahead of other open and proprietary models tested. That framing is important: the gap to the absolute frontier has narrowed, but the top closed tier still leads on the authors' own benchmarks.
Infrastructure contributions receive equal billing. The paper describes algorithm-system co-design for Kimi Delta Attention, balanced expert-parallel training, million-token agentic RL with persistent rollout states, and deployment optimizations needed to run a model at this scale.
Why open weights matter here
Releasing full Kimi K3 weights is not a symbolic gesture at 2.8T scale. Labs that can host and serve the model gain a testbed for agentic RL, long-context workflows, and vision-language pipelines without waiting on API tiers.
The tradeoff is operational: serving 104B activated parameters with MoE routing remains expensive even when total parameters are sparse. Enterprises evaluating Kimi K3 should budget for inference hardware and routing overhead, not just download bandwidth.
What we cannot verify from the paper alone
The arXiv submission is a tech report, not an independent audit. Moonshot's comparisons use internal suites alongside public benchmarks. Third-party replication on agentic and coding tasks will determine whether K3's open release shifts deployment decisions or remains a research milestone.
Still, the combination of scale, vision, million-token context, and public weights makes Kimi K3 one of the most significant open-model releases of mid-2026 — precisely because it targets the same long-horizon workloads closed labs are racing to monetize.
Sources
- arXiv — Kimi K3: Open Frontier Intelligence (July 27, 2026)
- Anthropic — Introducing Claude Opus 5 (July 24, 2026)
- OpenAI — Advancing the price-performance frontier with GPT-5.6 (July 30, 2026)