OPPD Distills Power-Distribution Sampling Into One Generation, Beating 64-Candidate Sampling
USC · hf · 2026-10-07
USC researchers propose on-policy power distillation (OPPD), which distills power-distribution sampling (raising each answer's probability to a power and renormalizing, sharpening toward the model's best answers) directly into the model so a single generation achieves what previously required many scored candidates.
The trained model generates candidates in an SMC-style loop; a frozen teacher's power distribution weights them, and the same weights drive the maximum-likelihood update.
Key results:
- Up to +23.0 on MATH500 and +27.3 on GSM8K over the untrained model at the same temperature, single generation.
- Beats published 64-candidate power sampling by 2.4 / 3.5 points with one generation.
- Versus GRPO at the same budget: +3.8 / +4.0 / +5.4 on MATH500 / GSM8K / AIME with no reference answers; complementary — OPPD after GRPO adds up to +9.3.
- Trained only on math, HumanEval still gains up to +5.3; gains hold across model families and sizes.
Code is open-sourced.
More from Research
- Anders Sandberg: co-writing papers with an LLM makes preregistering predictions painless — anderssandberg · 2026-10-07
- Prediction: zeroth-order optimization methods may eventually replace backprop — j_foerst · 2026-10-07
- Berkeley's Workhorse trains humanoid G1 for whole-body manipulation from human data only — pabbeel · 2026-10-07
- François Fleuret: my fancy new layer got crushed by good old algorithmic tricks — francoisfleuret · 2026-10-07
- Epoch AI launches Capabilities Index (ECI), a unified scale for comparing model intelligence over time — ricklamers · 2026-10-07
- ANVIL III Optimizer Claims 62% Pretraining Cost Cut at Frontier Scale, Beats Tuned Muon — kellerjordan0 · 2026-10-07