Self-Play Search Distillation boosts LLM math reasoning

A University of Edinburgh team introduced SPSD, which distills MuZero-style self-play search from board games into LLM reasoning chains, raising math benchmark scores from 24.1 to 36.6.

2026-09-29 ~ 2026-09-29 · 2 related posts