Ai2 Replaces Priority Scheduler: Teams Got 98% of Owed GPU Hours in 30-Day Test
allen_ai · x · 2026-10-09
Ai2's AI Infrastructure team details the design decisions and rollout lessons behind its GPU cluster scheduler. It frames scheduling around a four-layer pyramid of metrics: availability (how often hardware is healthy), occupancy (share of available time assigned to a workload), impact (how often the most valuable workloads get resources), and utilization (GPU capacity used over a workload's lifetime).
Context: Ai2 runs thousands of NVIDIA H100, B200 and B300 GPUs in clusters of 88 to 1024 GPUs, serving about 150 internal researchers across LLM/VLM training, robotics RL simulation and scientific agentic post-training. Demand runs 2-3x supply.
The team replaced a priority-based scheduler with GPU time budgets, hierarchical fair-share allocation and a time-slicing contract, turning "how much GPU time does each project deserve" from case-by-case operations into a transparent administrative budgeting process. In a 30-day test, teams received 98% of the GPU hours they were owed (accounting for actual demand), cluster occupancy stayed at 98%, and spare capacity went to interruptible work without drawing down a team's budget.
More from Infra
- Data centers could become SpaceX's biggest revenue stream, with $1T revenue on the horizon — Dr_Singularity · 2026-10-10
- New gTLD application for .lan shows why internal services need real registered domains — evilsocket · 2026-10-10
- Chart adds ByteDance Volcano Engine and Alibaba Cloud genAI procurement, classified by AI agent — Miles_Brundage · 2026-10-10
- 68GB Qwen MoE at 21 tok/s on RTX 3060 + 16GB RAM, bit-exact via new --moe-direct-io — zyxciss · 2026-10-10
- Krueger's New Paper: 'Hardwire' AI Models Into Chips Instead of Dismantling the Compute Supply Chain — DavidSKrueger · 2026-10-10
- TokenSpeed hits SOTA for Kimi K3 on AMD MI355X: 1.4x faster than ATOM at 216 tok/s — zhyncs42 · 2026-10-10