SGS paper: 7B model after 200 self-play rounds beats a 671B model pass@4
cephaloform · x · 2026-10-04
A Stanford team (Luke Bailey, Kaiyue Wen, Kefan Dong, Tatsunori Hashimoto, Tengyu Ma) proposes Self-Guided Self-Play (SGS) in arXiv:2604.20209.
- Problem: LLM self-play (a Conjecturer creates problems for a Solver, both improve together) should have unbounded learning, but existing methods plateau at scale. Over long runs the Conjecturer learns to hack its reward, collapsing to artificially complex problems that don't help the Solver.
- Method: SGS has the model take three roles — Solver, Conjecturer, and a Guide that scores synthetic problems by relevance to unsolved target problems plus cleanliness/naturalness, supervising against Conjecturer collapse. Core hypothesis: LLMs can judge whether a subproblem is useful for a goal.
- Results: On Lean4 formal theorem proving, SGS surpasses the asymptotic solve rate of the strongest RL baseline in under 80 rounds; a 7B model after 200 self-play rounds solves more problems than a 671B model pass@4.
Evaluated by training far longer than prior work and fitting scaling laws to cumulative solve-rate curves.
More from Research
- BF16 rounding breaks a conservation law, blowing up FlashAttention gradients late in training — HongyiWang10 · 2026-10-04
- Microsoft's ActiveSaddler adapts agent harness training scenarios, boosting Pass@1 by up to 7.5 points — dair_ai · 2026-10-04
- Anthropic's circuit tracing paper reverse-engineers how Claude 3.5 Haiku reasons internally — austinc3301 · 2026-10-04
- V-Rubrics: 50k visual samples split into 353k checkable criteria fix multimodal RL credit assignment — jiqizhixin · 2026-10-04
- Looped-DiT: 260M looped model beats 6.5x larger text-to-image rival with 4.9x less compute — Apprehensive_Sky892 · 2026-10-04
- Schmidhuber says he published the first concrete RSI algorithm back in 1987 — SchmidhuberAI · 2026-10-04