Skip to content
worth noting Hardware

Ai2 moved from job priorities to GPU time budgets

only one source so far

Ai2 replaced priority-based job scheduling with GPU time budgets, hierarchical fair-share allocation and time-slicing rules. Demand reaches two to three times the available capacity; according to the institute, the new system shifts allocation decisions to transparent budgeting.

Ai2 replaced job-priority scheduling with a system that combines GPU time budgets, hierarchical fair-share allocation and time-slicing rules. Managers can allocate GPU time proportionally among projects and researchers. According to the institute, this has shifted allocation decisions from individual operational disputes to transparent administrative budgeting.

The institute manages thousands of NVIDIA H100, B200 and B300 GPUs in clusters ranging from 88 to 1024 GPUs for approximately 150 internal researchers. Based on submitted jobs, demand regularly reaches two to three times the available capacity. The infrastructure is primarily used for large-scale distributed model training.

According to the institute, the previous system led to capacity being reserved through idle jobs and to priority inflation: eventually, 100 % of scheduled jobs had HIGH priority. Fixed GPU allocations to teams, meanwhile, left some hardware unused when a team had no experiments ready. The new approach therefore allocates a share of computing time instead of specific GPUs and allows management to make decisions in advance based on the expected value of the research. Details are available in the source article.

What changed

Why it matters

The case offers administrators of overloaded research clusters a concrete way to connect strategic priorities with capacity allocation. GPU time budgets determine project shares before jobs arrive, while fixed hardware allocations, according to the institute's experience, led to unused capacity during gaps between experiments.

Two audiences, two different impacts

What this means

01

For individuals

For internal researchers at Ai2, the way they obtain computing capacity is changing: instead of competing for higher job priorities, projects and individuals receive shares of GPU time through hierarchical budgeting.

More practical updates →
02

For a business

The change also affects the operations team. According to the institute, under the previous system, negotiating the termination of non-preemptible jobs on hardware requiring maintenance took up most of the time on-call engineers spent handling requests.

Processes More business impacts →
AI2 fair-share GPU time-slicing

Check the original

Event sources

only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
Hugging Face Blog primary source · first detected Impactful scheduling for GPU clusters