Self-Play Trio SPIRAL, SPICE, SPADE Push LLM Reasoning Up 5-10% Without Human Labels

青稞AI · wechat · 2026-09-21

A rundown of three self-play RL post-training works by Stanford PhD student Bo Liu, ahead of his livestream. SPIRAL (ICLR 2026) has models play multi-turn zero-sum games against their improving selves, generating an infinite curriculum with no human labels — training Qwen3-4B-Base on Kuhn Poker alone lifted math by 8.6% and general reasoning by 8.4%, beating SFT on 25,000 expert trajectories. SPICE (Meta FAIR) splits one model into a Challenger mining real corpora for problems and a Reasoner solving them, yielding +8.9% math and +9.8% reasoning. SPADE goes further: the model writes executable training environments as the designer, rewarded by the solver's gap between prompted and unprompted performance; a 30B model gained 5.3 points on average across eight benchmarks and +13.9 on ACEBench-Agent. The talk will also probe when self-improvement becomes genuinely recursive.

Related event: Stanford's Liu Bo to present SPIRAL/SPICE/SPADE self-play research(2 posts)→

Original post →

More from Research

Research channel →