Self-play on board games distills superhuman CoT, lifting Qwen3-4B math average from 24.1 to 36.6

PMinervini · x · 2026-09-29

New arXiv paper from the University of Edinburgh's Pasquale Minervini group introduces Self-Play Search Distillation (SPSD): MuZero-like networks self-play on board games, and the search records — preferred moves, alternatives, opponent replies, value estimates — are distilled into superhuman chains-of-thought for training LLM reasoning. On Qwen3-4B-Base, SPSD raises the mean over six math benchmarks from 24.1 to 36.6 and lifts held-out-game win rate from 15% to 45%, transferring to unseen math despite training only on game search data. The authors frame massive self-play over synthetic environments as an unbounded source of superhuman post-training data.

Related event: Self-Play Search Distillation boosts LLM math reasoning(2 posts)→

Original post →

More from Research

Research channel →