Ai2 rebuilt its GPU scheduler: median queue wait fell from 5 minutes to 24 seconds on H100 cluster

allen_ai · x · 2026-10-09

Ai2 shares the engineering behind its new GPU cluster scheduler. The team manages thousands of H100/B200/B300 GPUs in 88–1024-GPU clusters serving 150 researchers, with demand at 2–3x supply at any moment.

They frame scheduling as a pyramid: availability → occupancy → impact → utilization. They recently replaced a priority-based scheduler with a system using GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract, turning "how much GPU time each project deserves" from case-by-case ops into transparent budgeting. On the largest H100 cluster, median queue wait dropped from 5 minutes to 24 seconds.

Related event: Ai2 Rebuilt Its GPU Scheduler, Cutting Median Queue Time from 5 Minutes to 24 Seconds(3 posts)→

Original post →

More from Infra

Infra channel →