Pivot-SD: 200 questions beat diffusion RL by distilling only high-impact denoising pivots
kaist-ai · hf · 2026-10-05
Masked diffusion LMs suffer a credit-assignment problem: a few token commitments during denoising shape the whole response, yet most post-training ignores them. KAIST's Pivot-SD identifies these high-impact "pivots" via an information-gain metric, training successful pivots with cross-entropy and failed ones with targeted unlikelihood. Using just 200 questions with 4 rollouts each, it lifts LLaDA-8B-Instruct above full-sequence SFT and budget-matched diffusion RL baselines on math and code benchmarks.
More from Research
- Meta's NAVA-WAM pretrains robot action policies directly from action-free videos — meta · 2026-10-05
- gamfit: open-source Rust engine fits GAMs from a formula with REML-chosen smoothing — Sauers_ · 2026-10-05
- Fine-tuned Llama 3.1 8B for medical decisions hits 84% per-field accuracy, only 30-34% perfect rows — Forsaken_Cut8542 · 2026-10-05
- SJTU's LIFT adds force sensing to VLAs with zero force-labeled pretraining data — jiqizhixin · 2026-10-05
- LOOM stabilizes looped MoEs at 9-12 loops, beating standard MoE at iso-FLOP — SonglinYang4 · 2026-10-05
- Paper accepted at ACML 2026, but author can't afford to present it — Jealous_Key_4030 · 2026-10-05