AI2 has released Olmo-core 3, a redesigned open training infrastructure for mixture-of-experts (MoE) models that scales to the trillion-parameter range while maintaining computational efficiency. The framework switches from fully sharded data parallelism to distributed data parallelism, keeping experts resident on GPUs to improve throughput and reduce communication costs.

  • A benchmark on eight NVIDIA B300 GPUs showed a 47-billion-parameter MoE achieving 52,000 tokens per second per GPU, representing a 2.7x increase over the previous implementation.
  • Enabling MXFP8 precision on four NVIDIA B300 GPUs increased training throughput by approximately 21% and reduced peak active memory from 103 GiB to 95 GiB compared to BF16.
  • The system supports expert, pipeline, and optimizer parallelism to distribute model weights and training states across GPU clusters.
  • Benchmarks demonstrated scalability to a 1.2-trillion-parameter model with 858 TFLOP/s/GPU throughput and experimental configurations reaching 2.38 trillion parameters.

The open-source stack serves as the foundation for the next generation of Olmo models and allows researchers to train their own large MoEs with greater flexibility.