SPSD distills MuZero self-play search into LLM reasoning traces, math jumps 24.1→36.6
PMinervini · x · 2026-09-29
Researchers propose Self-Play Search Distillation (SPSD), asking whether self-play on board games can teach an LLM to reason:
- Method: train MuZero experts via self-play on board games, then distill their search into reasoning traces for LLMs.
- Results: reasoning transfers to unseen games and generalizes to math, improving scores from 24.1 to 36.6.
- Takeaway: planning ability from search-based RL can serve as a reasoning supervision signal beyond text-only distillation.
Related event: Self-Play Search Distillation boosts LLM math reasoning(2 posts)→
More from Research
- Google Cloud reproduces Olmo 3 7B pre-training on TPUs, matching Ai2 on held-out evals — allen_ai · 2026-09-29
- Researcher builds JevBench, a new 3,604-question evaluation benchmark pool — airesearch12 · 2026-09-29
- Microsoft Research unveils Project Quine, an AI research system combining biology world model with wet lab — erichorvitz · 2026-09-29
- Stanford's Kundaje Lab ports Illumina's PromoterAI to PyTorch, validates TERT promoter mutation scores — anshulkundaje · 2026-09-29
- Why Living Science Still Matters in the AI Era: Three Arguments — ChenhaoTan · 2026-09-29
- AI Revisits Cengiz et al.: 236 Minimum Wage Hikes Show No Detectable Job Loss — ChenhaoTan · 2026-09-29