Self-play on board games distills superhuman CoT, lifting Qwen3-4B math average from 24.1 to 36.6
PMinervini · x · 2026-09-29
New arXiv paper from the University of Edinburgh's Pasquale Minervini group introduces Self-Play Search Distillation (SPSD): MuZero-like networks self-play on board games, and the search records — preferred moves, alternatives, opponent replies, value estimates — are distilled into superhuman chains-of-thought for training LLM reasoning. On Qwen3-4B-Base, SPSD raises the mean over six math benchmarks from 24.1 to 36.6 and lifts held-out-game win rate from 15% to 45%, transferring to unseen math despite training only on game search data. The authors frame massive self-play over synthetic environments as an unbounded source of superhuman post-training data.
Related event: Self-Play Search Distillation boosts LLM math reasoning(2 posts)→
More from Research
- 6 best visual resources for learning Transformers, LLMs and diffusion — techNmak · 2026-09-29
- 180k typed tool-calling decisions released so tiny models can route tools — MaziyarPanahi · 2026-09-29
- Pruned CTC Cuts ASR Training Memory 5.1x, Enabling Native-LLM-Vocab Speech Recognition — X-LANCE · 2026-09-29
- SCOPD Distillation Keeps 92% of VLM Performance at 10% Visual Tokens — uoft · 2026-09-29
- AnswerMap: Training-Free Black-Box Spatial Rationale Hits 0.85 AUC vs 0.38 for Attention — Mohamed Eltahir · 2026-09-29
- Findings of ACL already separates papers from talks, so 'AI-authored' rules are moot, scholar argues — ipeirotis · 2026-09-29