Huawei has released the open-source weights for openPangu-2.0-Pro, a Mixture of Experts (MoE) large language model trained on Ascend hardware. The model features 505 billion total parameters with 18 billion activated parameters and supports a context length of 512k tokens.

  • Total pretraining data consists of 34T tokens.
  • Post-training utilizes unified SFT with slow and fast thinking capabilities.
  • Training includes multiple specialist RL training and on-policy distillation combining multiple RL specialists.

The model's technical details are documented in the openPangu-2.0 Tech Report.