
NVIDIA is investing $4 million(約6.4億円) over three years in CNCF to support open-source GPU infrastructure for Kubernetes, the platform already used by 66% of organizations running generative AI. The company contributed its GPU Dynamic Resource Allocation Driver as a vendor-neutral standard and backed the KAI Scheduler as a community project, aiming to move GPU management into open community standards rather than proprietary solutions. Only 7% of organizations currently deploy models daily, and NVIDIA frames this effort as essential to making continuous AI production as routine as any other workload.
Summaries like this, in your inbox every morning.
Sign up free →What happened
NVIDIA joined the CNCF Governing Board and committed $4 million(約6.4億円) over the next three years to fund CI and testing on real GPUs for CNCF projects. The company also contributed the GPU Dynamic Resource Allocation (DRA) Driver upstream to Kubernetes as a vendor-neutral reference implementation, and supported acceptance of the KAI Scheduler as a CNCF Sandbox project. The Kubernetes AI Conformance Program has grown from 18 to 31 certified platforms since launching.
Why it matters
Only 7% of organizations deploy AI models daily, and 47% only intermittently—a gap NVIDIA attributes to operational challenges at scale. The company argues that Kubernetes, already used by 66% of organizations hosting generative AI to manage inference workloads, lacks first-class GPU support; moving GPU management into open community standards rather than proprietary extensions allows teams to run AI reliably in continuous production without overprovisioning hardware or leaving accelerators idle.
What to watch
The Kubernetes AI Conformance Program's new v1.35 requirements, which cover agentic workflow support and in-place pod resizing for inference serving. NVIDIA's DRA Driver is in alpha for MIG device sharing and introduces ComputeDomains for safe GPU memory sharing across nodes via Multi-Node NVLink.
In July 2026, NVIDIA Senior Director Erin A. Boyd announced a deepened commitment to the Cloud Native Computing Foundation (CNCF), including a $4 million(約6.4億円) pledge over three years for CI and testing infrastructure. The move reflects growing operational challenges in running AI at scale on Kubernetes, the platform now used by 82% of container users in production and 66% of organizations hosting generative AI to manage at least some inference workloads.
Despite Kubernetes' dominance, adoption of continuous AI deployment remains low: only 7% of organizations deploy models daily, while 47% do so only intermittently. NVIDIA attributes much of this gap to orchestration inefficiencies. Historically, Kubernetes treated GPUs as static, indivisible resources—fine for early use cases but inadequate when organizations run large-scale training, latency-sensitive inference, and multi-tenant AI platforms on shared clusters simultaneously. This mismatch results in overprovisioned hardware, idle accelerators, and applications that fail in production despite passing tests.
To address this, NVIDIA contributed the GPU Dynamic Resource Allocation (DRA) Driver upstream to Kubernetes SIG-Node as a vendor-neutral reference implementation. The DRA Driver replaces static GPU assignment with real-time, on-demand allocation. It includes MIG device sharing (in alpha), allowing multiple workloads to share physical GPU resources, and introduces ComputeDomains, which enable GPUs to share memory safely and quickly across nodes using Multi-Node NVLink. In practical terms, a large multi-GPU job can now request exactly the accelerators it needs when it needs them, rather than holding a static reservation.
A second contribution, the KAI Scheduler, was accepted as a CNCF Sandbox project at KubeCon + CloudNativeCon Europe. KAI handles gang scheduling (ensuring dependent jobs are not evicted needlessly), hierarchical queues with Dominant Resource Fairness for multi-team clusters, and asynchronous binding that has run on clusters exceeding 10,000 GPUs. NVIDIA is the same engine NVIDIA relies on internally; by moving it to community governance, the company signals intent to let the entire ecosystem shape its direction rather than tie it to a product roadmap.
A third component, the Kubernetes AI Conformance Program, has expanded from 18 to 31 certified platforms since its launch months earlier, providing vendors and customers a way to verify that AI-ready infrastructure performs consistently across providers. Version 1.35 of the conformance program adds requirements for agentic workflow support and in-place pod resizing for inference serving. NVIDIA frames all three efforts as part of a conviction that foundational AI infrastructure should be a community asset. As the article quotes Jensen Huang: 'We want AI to diffuse into every industry and every country, every researcher, every student. And if everything is proprietary, it's hard to do research and it's hard to innovate on top. And so open source is fundamentally necessary.' The commitment reflects a historical observation: open, interoperable, community-governed platforms have consistently built larger ecosystems and moved faster than closed alternatives.
The article frames GPU orchestration on Kubernetes as a critical bottleneck limiting continuous AI deployment. NVIDIA cites data showing that only 7% of organizations deploy models daily, attributing this partly to operational difficulties: Kubernetes historically treated GPUs as static, indivisible resources, leading to overprovisioning and idle accelerators when teams run mixed workloads (training, latency-sensitive inference, multi-tenant platforms) on shared clusters. NVIDIA argues this is fundamentally an orchestration problem best solved through open, community-governed standards rather than vendor-specific extensions.
NVIDIA's new strategy reflects a shift in how the company engages with open-source infrastructure. Rather than shipping proprietary tools, it is contributing reference implementations (like the DRA Driver) directly into Kubernetes' standard governance, joining the CNCF Governing Board, and funding community testing infrastructure. The company frames this commitment as strategic but not purely altruistic: leadership quotes Jensen Huang saying that closed platforms make research and innovation harder, and the article asserts that historical precedent shows open, interoperable platforms (Linux, Kubernetes) build larger ecosystems and move faster than closed ones.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime
1 minute a day. The AI essentials.
200+ sources · Email / LINE / Slack