DeepSeek-AI presents DeepSeek-V3, a strong Mixture-of-Experts language model with 671B total parameters and 37B activated per token. The model adopts Multi-head Latent Attention and DeepSeekMoE architectures to achieve efficient inference and cost-effective training.

  • DeepSeek-V3 pioneers an auxiliary-loss-free strategy for load balancing to minimize performance degradation.
  • It employs a multi-token prediction training objective to enhance overall benchmark performance.
  • The team validates FP8 mixed precision training on an extremely large-scale model for the first time.
  • Training utilized 14.8 trillion tokens and cost only 2.788M H800 GPU hours, totaling approximately $5.576 million.
  • Evaluations show it outperforms other open-source models and matches leading closed-source models like GPT-4o and Claude-3.5-Sonnet.

The report emphasizes that DeepSeek-V3 achieves strong performance comparable to top-tier closed-source models while maintaining remarkably low training costs and stability.