Researchers introduce On-Policy Power Distillation (OPPD), a method that trains language models to generate high-quality answers in a single pass by mimicking the sharpened probability distribution of a teacher model. This approach eliminates the need for extensive candidate sampling during inference while significantly boosting performance on reasoning tasks.
- OPPD uses a sequential Monte Carlo sampler where a frozen teacher's power distribution weights candidates generated by the model being trained.
- The method increases single-generation accuracy by up to 23.0 points on MATH500 and 27.3 on GSM8K compared to the untrained model at the same temperature.
- A single generation from the distilled model outperforms published power sampling with 64 candidates, recovering 94 percent of the gain provided by 16 candidates in the untrained model.
- OPPD scores higher than GRPO on MATH500, GSM8K, and AIME using no reference answers, and can add up to 9.3 points when applied after GRPO.
- The technique also raises HumanEval accuracy by up to 5.3 points and works across different model families and sizes.
OPPD provides a parameter-efficient way to improve reasoning capabilities without changing the model's architecture or requiring search-based inference.