SPADE Paper: Single Model Self-Implements Environment Design and Agent Self-Play

_AndrewZhao · x · 2026-08-21

Natasha Jaques et al. release the SPADE paper, proposing recursive self-improvement via multi-agent RL and Unsupervised Environment Design (UED). The framework uses a single LLM to act as both an "Environment Designer" (building multi-turn RL training environments using the Gym step()/reset() API) and a "Reasoning Agent" that learns to solve them. The Designer optimizes task difficulty by maximizing the Agent's regret, achieving automatic scaling of training environments.

Related event: SPADE: Single-Model Self-Play Generates Ever-Harder RL Environments for Continued Self-Improvement(7 posts)→

Original post →

More from Research

Research channel →