Ai2's AI infrastructure team describes in a blog post how it overhauled scheduling for its GPU clusters, which run large distributed training workloads across thousands of NVIDIA H100, B200, and B300 GPUs. The clusters serve about 150 internal researchers, and demand for GPU time outstrips supply by a factor of two to three. The team's goal was to move from a system that rewarded gaming to one that transparently prioritizes high-impact research while maintaining full occupancy.
The previous priority-based scheduler allowed workloads to opt out of preemption, which led to predictable problems. Users parked no-op workloads to hold capacity, priority levels inflated until all workloads were marked HIGH, and on-call engineers spent most of their time negotiating shutdowns of non-preemptable jobs on hosts needing maintenance. The team recognized this as a tragedy of the commons: individuals maximizing their own outcomes degraded the shared resource for everyone.
Their solution was to stop trying to solve the scheduling puzzle directly. Instead of assigning teams monopolies over GPUs, which caused idle capacity due to research seasonality, they allocated portions of GPU time. This turns the question of which research deserves resources into an administrative budgeting decision made in advance, rather than a case-by-case operational fight. The new system combines GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract, allowing leadership to act like investors funding research efforts based on likely impact.