Moonshot has released the full details, weights, and supporting infrastructure for Kimi K3, a 2.8T-parameter Mixture-of-Experts model with approximately 104B active parameters per token. The release includes technical reports detailing its hybrid long-context stack and new tools like MoonEP, FlashKDA, and AgentEnv.

  • Kimi K3 scales across length, depth, and width using Kimi Delta Attention, Gated MLA, and sparse LatentMoE, trained via multi-teacher on-policy distillation.
  • Infrastructure components are central to the release, with deployment support from NVIDIA for Dynamo and Red Hat AI providing FP8-Block Hopper-tuned checkpoints.
  • Cost analysis indicates significant hardware requirements, with minimum configurations needing 8x MI355X GPUs and production serving potentially requiring 64+ GPUs in a high-bandwidth domain.
  • The open-weight release has prompted rapid adoption by providers like Perplexity, Baseten, and Together, who are offering hosted inference options due to the high cost of self-hosting.

The release underscores that open weights for frontier models now require substantial infrastructure investment, shifting access towards hosted offerings rather than simple self-deployment.