Researchers introduce On-Policy Power Distillation (OPPD), a method that trains language models to generate high-quality answers in a single pass by mimicking the sharpened probability distribution of a teacher model. This approach eliminates the need for extensive candidate sampling during inference while significantly boosting performance on reasoning tasks.

  • OPPD uses a sequential Monte Carlo sampler where a frozen teacher's power distribution weights candidates generated by the model being trained.
  • The method increases single-generation accuracy by up to 23.0 points on MATH500 and 27.3 on GSM8K compared to the untrained model at the same temperature.
  • A single generation from the distilled model outperforms published power sampling with 64 candidates, recovering 94 percent of the gain provided by 16 candidates in the untrained model.
  • OPPD scores higher than GRPO on MATH500, GSM8K, and AIME using no reference answers, and can add up to 9.3 points when applied after GRPO.
  • The technique also raises HumanEval accuracy by up to 5.3 points and works across different model families and sizes.

OPPD provides a parameter-efficient way to improve reasoning capabilities without changing the model's architecture or requiring search-based inference.