Moonshot AI has introduced Kimi K3, an open-source Mixture-of-Experts model featuring 2.8 trillion total parameters and 104 billion activated parameters per token. The model supports native vision capabilities and a 1-million-token context window, leveraging Kimi Delta Attention and Stable LatentMoE to achieve approximately 2.5x scaling efficiency over its predecessor, Kimi K2.

  • Kimi K3 utilizes Kimi Delta Attention and Attention Residuals to improve information flow across sequence length and model depth.
  • Stable LatentMoE effectively activates 16 of 896 routed experts per token.
  • Post-training includes reinforcement learning across general, agentic, and coding domains with multiple reasoning-effort levels.
  • The model achieves frontier-level performance in long-horizon coding, agentic, knowledge, reasoning, and vision tasks.
  • While trailing Claude Fable 5 and GPT-5.6 Sol, it outperforms other evaluated open and proprietary models.

Moonshot AI is releasing the full Kimi K3 model weights to facilitate future research and accelerate the broader deployment of frontier intelligence.