SPADE: Self-Play Where Agents Design Ever-Harder RL Environments, a Step Toward RSI

_AndrewZhao · x · 2026-08-21

SPADE♠ (Self-Play for Agentic Environment Scaling) has one model self-play the Environment Designer and the Reasoning Agent: the designer continuously writes executable, agentic RL environments that get harder as the agent improves, automating environment scaling. The authors argue it's a key step toward RSI (recursive self-improvement) — AI deciding what to train and verify. More settings covering terminal tasks are in the works.

Related event: SPADE: Single-Model Self-Play Designs Executable Environments for Recursive Self-Improvement(5 posts)→

Original post →

More from AGI Musings

AGI Musings channel →