Researchers introduce On-Policy Power Distillation (OPPD), a method that trains models to generate high-quality answers in a single pass by mimicking the sharpened probability distribution of a teacher model. This approach eliminates the need for extensive candidate sampling during inference while significantly boosting performance on mathematical and coding benchmarks.
- OPPD uses a sequential Monte Carlo sampler where a frozen teacher's power distribution weights candidates generated by the student model for maximum-likelihood updates.
- The method increases single-generation accuracy by up to 23.0 points on MATH500 and 27.3 points on GSM8K compared to the untrained baseline at the same temperature.
- A single generation from the distilled model outperforms power sampling with 64 candidates, recovering 94 percent of the gain provided by 16 candidates in the untrained model.
- OPPD scores higher than GRPO trained with verified rewards on MATH500, GSM8K, and AIME without using reference answers, and can add up to 9.3 points when applied after GRPO.
- The technique raises HumanEval accuracy by up to 5.3 points and maintains gains across different model families and sizes.
This method allows models to achieve reasoning improvements typically requiring heavy sampling or additional reference data through direct parameter updates.