Ai2 moved from job priorities to GPU time budgets
Ai2 replaced priority-based job scheduling with GPU time budgets, hierarchical fair-share allocation and time-slicing rules. Demand reaches two to three times the available capacity; according to the institute, the new system shifts allocation decisions to transparent budgeting.
Ai2 replaced job-priority scheduling with a system that combines GPU time budgets, hierarchical fair-share allocation and time-slicing rules. Managers can allocate GPU time proportionally among projects and researchers. According to the institute, this has shifted allocation decisions from individual operational disputes to transparent administrative budgeting.
The institute manages thousands of NVIDIA H100, B200 and B300 GPUs in clusters ranging from 88 to 1024 GPUs for approximately 150 internal researchers. Based on submitted jobs, demand regularly reaches two to three times the available capacity. The infrastructure is primarily used for large-scale distributed model training.
According to the institute, the previous system led to capacity being reserved through idle jobs and to priority inflation: eventually, 100 % of scheduled jobs had HIGH priority. Fixed GPU allocations to teams, meanwhile, left some hardware unused when a team had no experiments ready. The new approach therefore allocates a share of computing time instead of specific GPUs and allows management to make decisions in advance based on the expected value of the research. Details are available in the source article.
Why it matters
The case offers administrators of overloaded research clusters a concrete way to connect strategic priorities with capacity allocation. GPU time budgets determine project shares before jobs arrive, while fixed hardware allocations, according to the institute's experience, led to unused capacity during gaps between experiments.
Two audiences, two different impacts
What this means
For individuals
For internal researchers at Ai2, the way they obtain computing capacity is changing: instead of competing for higher job priorities, projects and individuals receive shares of GPU time through hierarchical budgeting.
More practical updates →For a business
The change also affects the operations team. According to the institute, under the previous system, negotiating the termination of non-preemptible jobs on hardware requiring maintenance took up most of the time on-call engineers spent handling requests.
Processes More business impacts →Check the original
Event sources
only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.