Multi-turn agent RL training at scale on HF Hub: 9,523 sandboxes in 14h, zero crashes

vanstriendaniel · x · 2026-09-16

First multi-turn agent harness RL training with sandboxing running at scale on the Hugging Face Hub: 9,523 sandboxes spawned in 14 hours, full-parameter training of Qwen3-Coder-30B-A3B (30B MoE) with FSDP2 + expert parallel, all on HF Jobs with no Slurm and zero crashes.

Original post →

More from coding & agent

coding & agent channel →