DiPOD stabilizes diffusion LM post-training, lifting Sudoku accuracy from 22% to 97% with a one-line change
berkeley_ai · x · 2026-10-10
- At the COLM workshop on Non-Autoregressive Language Models, Haozhe Jiang presented DiPOD (Diffusion Policy Optimization without Drifting Apart), work from the Berkeley AI team.
- The core claim: diffusion LMs underperform mainly because post-training is unstable, and DiPOD acts as a stable "tripod" for diffusion model post-training.
- It boosts accuracy across reasoning tasks, with Sudoku jumping from 22% to 97% — achieved with a one-line code change.
- Talk time 15:25–15:40 at Union Square 22, Fourth Floor.
More from Research
- Coevolved robot communication transfers poorly to 3D: 1 success in 30 seeds — uv-mex · 2026-10-10
- Transferring co-evolved robot communication from 2D to 3D physics simulation — uv-mex · 2026-10-10
- SGS tweak to RL resets lets sim-trained robots mesh gears at 94% zero-shot — abhishekunique7 · 2026-10-10
- SQUISH: 355k rectangular packings improve 15 square-packing records with near-zero compute — ctjlewis · 2026-10-10
- Lessons from planning rater studies in AI: a five-step recipe — davidstutz92 · 2026-10-10
- Regents Labs launches Paper Pro Daily, an AI paper-a-day digest built with ChatGPT Pro — seanwbren · 2026-10-10