AI2 has released Olmo-core 3, a redesigned open training infrastructure for mixture-of-experts (MoE) models that scales to the trillion-parameter range while maintaining computational efficiency. The framework switches from fully sharded data parallelism to distributed data parallelism, keeping experts resident on GPUs to improve throughput and reduce communication costs.
- A benchmark on eight NVIDIA B300 GPUs showed a 47-billion-parameter MoE achieving 52,000 tokens per second per GPU, representing a 2.7x increase over the previous implementation.
- Enabling MXFP8 precision on four NVIDIA B300 GPUs increased training throughput by approximately 21% and reduced peak active memory from 103 GiB to 95 GiB compared to BF16.
- The system supports expert, pipeline, and optimizer parallelism to distribute model weights and training states across GPU clusters.
- Benchmarks demonstrated scalability to a 1.2-trillion-parameter model with 858 TFLOP/s/GPU throughput and experimental configurations reaching 2.38 trillion parameters.
The open-source stack serves as the foundation for the next generation of Olmo models and allows researchers to train their own large MoEs with greater flexibility.