DeepSeek-AI presents DeepSeek-V3, a strong Mixture-of-Experts language model with 671B total parameters and 37B activated per token. The model adopts Multi-head Latent Attention and DeepSeekMoE architectures to achieve efficient inference and cost-effective training.
- DeepSeek-V3 pioneers an auxiliary-loss-free strategy for load balancing to minimize performance degradation.
- It employs a multi-token prediction training objective to enhance overall benchmark performance.
- The team validates FP8 mixed precision training on an extremely large-scale model for the first time.
- Training utilized 14.8 trillion tokens and cost only 2.788M H800 GPU hours, totaling approximately $5.576 million.
- Evaluations show it outperforms other open-source models and matches leading closed-source models like GPT-4o and Claude-3.5-Sonnet.
The report emphasizes that DeepSeek-V3 achieves strong performance comparable to top-tier closed-source models while maintaining remarkably low training costs and stability.