Ai2 Replaces Priority Scheduler: Teams Got 98% of Owed GPU Hours in 30-Day Test

allen_ai · x · 2026-10-09

Ai2's AI Infrastructure team details the design decisions and rollout lessons behind its GPU cluster scheduler. It frames scheduling around a four-layer pyramid of metrics: availability (how often hardware is healthy), occupancy (share of available time assigned to a workload), impact (how often the most valuable workloads get resources), and utilization (GPU capacity used over a workload's lifetime).

Context: Ai2 runs thousands of NVIDIA H100, B200 and B300 GPUs in clusters of 88 to 1024 GPUs, serving about 150 internal researchers across LLM/VLM training, robotics RL simulation and scientific agentic post-training. Demand runs 2-3x supply.

The team replaced a priority-based scheduler with GPU time budgets, hierarchical fair-share allocation and a time-slicing contract, turning "how much GPU time does each project deserve" from case-by-case operations into a transparent administrative budgeting process. In a 30-day test, teams received 98% of the GPU hours they were owed (accounting for actual demand), cluster occupancy stayed at 98%, and spare capacity went to interruptible work without drawing down a team's budget.

Related event: Ai2 Rebuilt Its GPU Scheduler, Cutting Median Queue Time from 5 Minutes to 24 Seconds(3 posts)→

Original post →

More from Infra

Infra channel →