SortedRL: Microsoft Research tackles 70-74% GPU idle time in LLM reinforcement learning
burkov · x · 2026-09-30
Scaling RL for LLM reasoning is bottlenecked by length variance: batch updates must wait for the longest response, leaving GPUs idle 70–74% of generation cycles. Microsoft Research's SortedRL is an online, length-aware scheduling system that speeds up training and improves learning efficiency without destabilizing the RL process.
More from Infra
- DeepSeek open-sources full Ascend software stack, key benchmarks near hardware limits — lxfater · 2026-09-30
- SAXON Q: A room-temperature diamond-chip quantum computer that plugs into a wall outlet — TinfoilTricorn · 2026-09-30
- Where should neolabs get GPUs? Insider votes put SF Compute ahead of CoreWeave — FinanceYF5 · 2026-09-30
- Google donates Agent Substrate to CNCF as runtime layer for large-scale agents — rakyll · 2026-09-30
- South Korea expects record-high tax revenue amid AI chip boom — Polymarket · 2026-09-30
- Government weather model WRF ported to GPUs, running an order of magnitude faster with 250m fog forecasts — Scobleizer · 2026-09-30