SPSD distills MuZero self-play search into LLM reasoning traces, math jumps 24.1→36.6

PMinervini · x · 2026-09-29

Researchers propose Self-Play Search Distillation (SPSD), asking whether self-play on board games can teach an LLM to reason:

Related event: Self-Play Search Distillation boosts LLM math reasoning(2 posts)→

Original post →

More from Research

Research channel →