One sandbox per rollout: how frontier labs run RL for agents in 2026
SergioPaniego · x · 2026-09-11
A companion blog to Class 4 of the Training Agents series (live training of coding agents in real environments with TRL and OpenEnv), and the third recap of frontier model reports:
- The problem: agents run code, browse, and edit files over many turns, so training them requires a place where all of that happens, plus machinery keeping thousands of those places running simultaneously.
- Content: after covering how frontier labs use distillation and outcome-based training in earlier recaps, the author maps the sandbox/environment architectures labs built to run agent RL at scale in 2026.
More from Infra
- Qdrant lines up three free community events with 4-hour vector tech stream — qdrant_engine · 2026-09-11
- B300 spot prices hit $2.2M per unit in China, 3x premium pushes domestic chips into the推理 sweet spot — aigclink · 2026-09-11
- Self-hosted Qwen 27B on RunPod hits only 17 tok/s generation, making OpenAI API hard to beat on cost per job — yeah280 · 2026-09-11
- antirez: DeepSeek's shared KV cache architecture enables fast big prefills on low-memory local setups — antirez · 2026-09-11
- vLLM details MiniMax M3 optimization on AMD MI355X: 4.45x per-GPU throughput gains — vllm_project · 2026-09-11
- antirez breaks down DwarfStart's three-tier prefill strategy on local hardware — antirez · 2026-09-11