SortedRL: Microsoft Research tackles 70-74% GPU idle time in LLM reinforcement learning

burkov · x · 2026-09-30

Scaling RL for LLM reasoning is bottlenecked by length variance: batch updates must wait for the longest response, leaving GPUs idle 70–74% of generation cycles. Microsoft Research's SortedRL is an online, length-aware scheduling system that speeds up training and improves learning efficiency without destabilizing the RL process.

Original post →

More from Infra

Infra channel →