AIToday
AI Business & IndustryHugging Face BlogPublished: Oct 10, 2026, 01:00 JST

Ai2's new GPU scheduler delivers 98% of owed hours

Ai2's new GPU scheduler delivers 98% of owed hours

3 Key Points

  1. What happened

    Ai2 replaced its priority-based GPU scheduler with GPU time budgets and fair-share allocation, delivering 98% of owed GPU hours over a 30-day test and cutting human-in-the-loop repairs by 74%.

  2. Why it matters

    Clearly defined budgets and time-slicing let teams reclaim unused capacity later, which one researcher described as feeling like 30% more compute, and prevent costly workarounds like GPU squatting.

WHO IT HITSThis affects AI research teams and their managers who allocate scarce GPU capacity among competing projects; they may benefit from clearer, budget-based ways to prioritize work without leaving expensive hardware idle.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Ai2 runs thousands of NVIDIA H100, B200, and B300 GPUs across clusters ranging from 88 to 1024 GPUs, serving about 150 internal researchers. Demand for GPU time has consistently exceeded supply, with outstanding requests for 2-3x more GPUs than are available. The old priority-based scheduler let workloads opt out of preemptability, which led to GPU 'squatting' and priority inflation, where eventually 100% of scheduled workloads used HIGH priority. On-call engineers spent most of their ticket time negotiating shutdowns of non-preemptable workloads on hosts needing maintenance.

The new budget-based system assigns each project a guaranteed share of GPU time, such as Project A1 having a 35% claim on total capacity. A hierarchical fair-share scheduler tracks occupancy over a 7-day lookback window and sorts workloads accordingly. Workloads must declare a minimum runtime in exchange for preemption protection; minimum runtime can be set to zero for free but preemptible use. The rollout began at the end of July, cluster by cluster.

Simulations using historical data and constructed scenarios predicted debug workload wait times would fall from about 6 hours to 5 minutes, and real outcomes overperformed, with p90 waits dropping from 2 hours to 30 seconds. Occupancy held steady at 98% before and after, with 18% of delivered GPU time unallocated. The learning curve was steeper than assumed, with live explanatory sessions proving more effective than documentation alone.

FAQ
How much did queue wait times improve for small debug workloads?
Under the new scheduler, p90 debug workload wait time fell from 2 hours to 30 seconds in production, compared to a simulated prediction of 6 hours to 5 minutes from hand-crafted scenarios.
What is the maximum protected runtime for a workload?
Workloads must declare a minimum runtime, and the maximum value allowed for minimum runtime is 8 hours; after that, a workload may be preempted if it exceeds its allocation.
What tradeoff did interactive sessions face?
Researchers using interactive sessions could previously hold them for up to a week, but with time-slicing they became subject to the 8-hour cap on protected runtime, after which a session could be preempted.
Hugging Face BlogRead Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articlePostman runs Agent Mode for 40 million developers on Amazon Bedrock