Ai2 rebuilt its GPU scheduler: median queue wait fell from 5 minutes to 24 seconds on H100 cluster
allen_ai · x · 2026-10-09
Ai2 shares the engineering behind its new GPU cluster scheduler. The team manages thousands of H100/B200/B300 GPUs in 88–1024-GPU clusters serving 150 researchers, with demand at 2–3x supply at any moment.
They frame scheduling as a pyramid: availability → occupancy → impact → utilization. They recently replaced a priority-based scheduler with a system using GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract, turning "how much GPU time each project deserves" from case-by-case ops into transparent budgeting. On the largest H100 cluster, median queue wait dropped from 5 minutes to 24 seconds.
More from Infra
- Amazon drops data center NDAs as community backlash spurs hundreds of moratoriums — TechCrunch AI · 2026-10-10
- Amazon drops data center NDAs, and AI agents want your credit card — TechCrunch AI · 2026-10-10
- SkyPilot founder: hoarded idle GPUs waste $20M+ a year, and the AI compute layer should be open — skypilot_org · 2026-10-10
- uv binary shrinks over 40% since July, saving nearly 4 PB of mirror bandwidth monthly — charliermarsh · 2026-10-10
- Firmus' $30B IPO collapses: only 46MW operational, valuation tripled in two months — kevinsxu · 2026-10-10
- Hyperscaler bonds now a notable share of net new Treasury borrowing — matt_slotnick · 2026-10-10