RL infra detail: re-prefill over PipelineRL's KV cache reuse, batch size for GPU utilization
stochasticchasm · x · 2026-09-22
A technical commentary on an RL training setup notes two interesting engineering choices: instead of PipelineRL's approach of keeping the old KV cache, the team opts to re-prefill; and large batch sizes are motivated as helping GPU utilization more than any training effects. The author also speculates what "behaviors" means — possibly optimized code/performance measured via relative ranking — and hopes the released environments provide more detail.
Related event: Debate Erupts Over Million-Scale GRPO RL Training Details(3 posts)→
More from Infra
- PyroDash cuts inference cost 96% by having a 4B model call the big one only when needed — jiqizhixin · 2026-09-22
- Underdog's Husky Inference Engine Claims 4.5x Speedup Over MLX, 730 tok/s on MacBook — jimmykoppel · 2026-09-22
- Full-Parameter RL on TPUs: peano_ai Runs 310B MiMo-V2.6 Across 1,000+ TPUs — simonguozirui · 2026-09-22
- Cloudflare Python Workers go generally available after two-year preview — Simon Willison · 2026-09-22
- Fighting AI crawler traffic: beyond Turnstile, Cloudflare's AI Labyrinth as an option — fforres · 2026-09-22
- Software moats won't survive RSI — ML infra's value is demand aggregation, says cHHillee — PatrickToulme · 2026-09-22