Multi-turn agent RL training at scale on HF Hub: 9,523 sandboxes in 14h, zero crashes
vanstriendaniel · x · 2026-09-16
First multi-turn agent harness RL training with sandboxing running at scale on the Hugging Face Hub: 9,523 sandboxes spawned in 14 hours, full-parameter training of Qwen3-Coder-30B-A3B (30B MoE) with FSDP2 + expert parallel, all on HF Jobs with no Slurm and zero crashes.
More from coding & agent
- Graph Engineering: From Monolithic Agent Loops to State-Managed Workflows — Pavan_Belagatti · 2026-09-16
- How Alertly's Founder Grew to 22k Users and $1M+ Revenue, Plus a New MCP Server — kylegawley · 2026-09-16
- Benzi coding agent hits 78.2% SWE-bench reading far less code than Claude Code — DonkeyTheKing · 2026-09-16
- Mystery stealth model hits OpenRouter, built for coding and agentic workflows — kevinkern · 2026-09-16
- A Week of Nonstop Flights, Saved by Codex Remote and ChatGPT Work — reach_vb · 2026-09-16
- AI coding output up 25% but code duplication jumped 81%, hidden costs emerge — rseroter · 2026-09-16