Huawei has released the open-source weights for openPangu-2.0-Pro, a Mixture of Experts (MoE) large language model trained on Ascend hardware. The model features 505 billion total parameters with 18 billion activated parameters and supports a context length of 512k tokens.
- Total pretraining data consists of 34T tokens.
- Post-training utilizes unified SFT with slow and fast thinking capabilities.
- Training includes multiple specialist RL training and on-policy distillation combining multiple RL specialists.
The model's technical details are documented in the openPangu-2.0 Tech Report.