Reflection has announced Beam, a text-only Mixture-of-Experts model with 501 billion total parameters and 23 billion active parameters, designed for coding, agentic, and scientific tasks. The model was trained from scratch using 23.8 trillion pretraining tokens and is scheduled to release full weights under the Apache 2.0 license this month.
- Beam utilizes a stable RL/OPD run on 10,000 GB300s with over 100 million rollouts across approximately one million tasks.
- Reflection claims an 80.9 score on SWE-bench Verified and reports inference efficiency three to four times that of GLM 5.2.
- The training data included an OCR pipeline processing hundreds of millions of PDFs, with pretraining and RL each taking about four weeks.
- Independent analysis suggests Beam is among the most token-efficient open models for its intelligence level, though it trails some Chinese counterparts like DeepSeek V4.1 Flash on certain benchmarks.
The release marks Reflection's arrival as a functional neolab and provides more options in the US-trained open model market.