Self-Play Search Distillation boosts LLM math reasoning
A University of Edinburgh team introduced SPSD, which distills MuZero-style self-play search from board games into LLM reasoning chains, raising math benchmark scores from 24.1 to 36.6.
2026-09-29 ~ 2026-09-29 · 2 related posts
- SPSD distills MuZero self-play search into LLM reasoning traces, math jumps 24.1→36.6 — PMinervini · 2026-09-29
- Self-play on board games distills superhuman CoT, lifting Qwen3-4B math average from 24.1 to 36.6 — PMinervini · 2026-09-29