SPADE: Self-Play in Adaptive Synthetic Executable Environments
Bo Liu · hf · 2026-08-20
SPADE is a self-play reinforcement learning framework where an LLM designs adaptive executable training environments and learns to solve them. It improves reasoning and tool-use performance through regret-based environment targeting.
More from Research
- Stanford Framework Boosts DeepSeek Past Claude at 1/11th Cost — FuSheng_0306 · 2026-08-20
- 3DGS-Uncertainty: Predictive Photometric Uncertainty for Gaussian Splatting — rsasaki0109 · 2026-08-20
- SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation — FrameXAI · 2026-08-20
- Zetta ζ: Closed-Loop Embodied Harness Evolves Critics and Recovery Skills Online — Xin Ding · 2026-08-20
- SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset for Deformable-Object Manipulation — TuojingAI · 2026-08-20
- Decision-Metric Alignment in Latent World Models for MPC Planning — simple-world-lab · 2026-08-20