SLAI introduces T-Rex, a full-parameter post-training framework for the DeepSeek-V4 model family on Ascend SuperPOD infrastructure. The system addresses memory pressure and communication overhead in trillion-parameter MoE models through hierarchical optimization.
- Achieves 34.22% Model FLOPs Utilization (MFU), a 2.93x improvement over the open-source baseline.
- Establishes a CPT and SFT workflow for Operations Research tasks using solver-verified synthetic data.
- The specialized DeepSeek-V4-Flash model reaches a 71.81% zero-shot Pass@1 score, outperforming GPT-5.4-Mini by 3.98 percentage points.
This work demonstrates an efficient pathway for post-training frontier models on Ascend hardware and creating domain-specialized capabilities for complex reasoning.