OPPD Distills Power-Distribution Sampling Into One Generation, Beating 64-Candidate Sampling

USC · hf · 2026-10-07

USC researchers propose on-policy power distillation (OPPD), which distills power-distribution sampling (raising each answer's probability to a power and renormalizing, sharpening toward the model's best answers) directly into the model so a single generation achieves what previously required many scored candidates.

The trained model generates candidates in an SMC-style loop; a frozen teacher's power distribution weights them, and the same weights drive the maximum-likelihood update.

Key results:

Code is open-sourced.

Original post →

More from Research

Research channel →