A Wild Async RL Config: 30 Steps x 25K Rollouts Per Step at Parallelism 4
willcb · x · 2026-09-22
A practitioner shared a notably aggressive RL training configuration: 30 training steps, 25,000 rollouts per step, with async parallelism of 4, calling it "a WILD config."
Configs at this rollout scale are typical of agent-style RL runs that need massive environment interaction, highlighting how async sampling infrastructure is being pushed for RL training.
More from Infra
- Fighting AI crawler traffic: beyond Turnstile, Cloudflare's AI Labyrinth as an option — fforres · 2026-09-22
- Software moats won't survive RSI — ML infra's value is demand aggregation, says cHHillee — PatrickToulme · 2026-09-22
- Raspberry Pi locks devices to original RAM size, blocking aftermarket memory upgrades — ngxson · 2026-09-22
- fal's H3 Max generates 5 seconds of frontier-quality video in just 3 seconds — gorkem · 2026-09-22
- NVIDIA's EPD Disaggregation Cuts Multimodal TTFT Up to 5x, E2E Latency 7x — dl_weekly · 2026-09-22
- Running MiniMax H3 locally on a 16GB Mac: 8-10s clips in 15-20 minutes — coberholzer · 2026-09-22