Pivot-SD: 200 questions beat diffusion RL by distilling only high-impact denoising pivots

kaist-ai · hf · 2026-10-05

Masked diffusion LMs suffer a credit-assignment problem: a few token commitments during denoising shape the whole response, yet most post-training ignores them. KAIST's Pivot-SD identifies these high-impact "pivots" via an information-gain metric, training successful pivots with cross-entropy and failed ones with targeted unlikelihood. Using just 200 questions with 4 rollouts each, it lifts LLaDA-8B-Instruct above full-sequence SFT and budget-matched diffusion RL baselines on math and code benchmarks.

Original post →

More from Research

Research channel →