Self-Play Trio SPIRAL, SPICE, SPADE Push LLM Reasoning Up 5-10% Without Human Labels
青稞AI · wechat · 2026-09-21
A rundown of three self-play RL post-training works by Stanford PhD student Bo Liu, ahead of his livestream. SPIRAL (ICLR 2026) has models play multi-turn zero-sum games against their improving selves, generating an infinite curriculum with no human labels — training Qwen3-4B-Base on Kuhn Poker alone lifted math by 8.6% and general reasoning by 8.4%, beating SFT on 25,000 expert trajectories. SPICE (Meta FAIR) splits one model into a Challenger mining real corpora for problems and a Reasoner solving them, yielding +8.9% math and +9.8% reasoning. SPADE goes further: the model writes executable training environments as the designer, rewarded by the solver's gap between prompted and unprompted performance; a 30B model gained 5.3 points on average across eight benchmarks and +13.9 on ACEBench-Agent. The talk will also probe when self-improvement becomes genuinely recursive.
Related event: Stanford's Liu Bo to present SPIRAL/SPICE/SPADE self-play research(2 posts)→
More from Research
- GaME (CVPR 2026): Gaussian mapping that forgets stale geometry as robot scenes change — lucacarlone1 · 2026-09-21
- YOCO back in spotlight: blog breaks down cross-layer KV sharing in DeepSeek-V4.1-Flash and Gemma 4 — donglixp · 2026-09-21
- Dead Human Brain Tissue Controls Robot: Hybrots Are 20 Years Old, Argue Critics — ryunuck · 2026-09-21
- New paper: simple difference-of-means vectors detect reward hacking from LLM internals before it happens — burny_tech · 2026-09-21
- What AI means for mathematicians: seven predictions extrapolated from software — cgarciae88 · 2026-09-21
- Qwen releases RecreationBench: 250 tasks testing agents that rebuild real apps — burny_tech · 2026-09-21