SPADE: Self-Play in Adaptive Synthetic Executable Environments

Bo Liu · hf · 2026-08-20

SPADE is a self-play reinforcement learning framework where an LLM designs adaptive executable training environments and learns to solve them. It improves reasoning and tool-use performance through regret-based environment targeting.

Original post →

More from Research

Research channel →