
AMD has made substantial progress in narrowing the software and hardware gap with Nvidia, securing major wins including a 2GW Anthropic deployment and Microsoft's renewed commitment to the MI455X Helios system. The MI455X represents a silicon engineering lead, shipping 2nm compute at 3,470mm² per package with advanced on-package integration. However, AMD's internal capacity planning poses a critical risk: the lack of stable GPU clusters for software development and continuous integration testing is blocking the software team from reaching its Advancing AI 2026 targets (such as 90% parity with CUDA on vLLM gating tests), and the timeline has already slipped to October 2026. Unless AMD's leadership reprioritizes GPU cluster stability and capacity, this constraint could undermine the company's ability to capitalize on its silicon advantages.
Summaries like this, in your inbox every morning.
Sign up free →What happened
AMD has upgraded its assessment of its ability to compete with Nvidia in AI accelerators, moving from "non-zero" to "great chance of success." Anthropic has announced a 2GW deployment of AMD chips, and Microsoft has committed to deploying the MI455X Helios system after previously skipping AMD generations. AMD's MI455X chip uses 2nm process technology and reaches 3,470mm² of logic silicon in a single package, outpacing competitors still on 3nm.
Why it matters
AMD's open-sourced compiler and kernels position it better for the emerging era of AI coding agents, and a novel equity-rebate structure with OpenAI and Meta makes AMD hardware cost-effectively negative when combined with these deals. However, AMD faces a critical internal bottleneck: stable GPU clusters for software development and continuous integration testing are severely constrained, blocking progress on vLLM gating tests (currently 0% parity with Nvidia's ConnectX CI for Kubernetes inferencing) and preventing AMD engineers from reaching the target of 90% parity with CUDA vLLM gating by Advancing AI 2026.
What to watch
AMD's internal vLLM team has been pulled away from stable clusters due to leadership's capacity-crunch decisions, pushing the planned open-source vLLM and SGLang CI target from Advancing AI 2026 to October 2026 (with an effort to shift back to August/September 2026). The Helios rack production ramp is slow due to non-cableless tray design and backplane reliability challenges requiring over 550 Broadcom ethernet retimers per rack. AMD's ability to solve these two risks—stable internal GPU capacity and rack production reliability—will determine whether the competitive momentum translates to market share gains.
AMD has substantially upgraded its internal assessment of its ability to compete with Nvidia in AI accelerators. In an earlier article, the authors assigned AMD a 0% chance; in a second article months later, they upgraded to "non-zero." Now, after observing AMD's software stack and leadership execution over the past year, they have moved to assessing AMD as having a "great chance of success"—contingent on solving two major risks around supply-chain/production and internal GPU capacity.
The competitive position is strengthening visibly on the customer front. Anthropic has publicly announced a 2GW deployment of AMD chips, and Anthropic Head of Compute Tom Brown has cited using Claude with a "/goal" instruction to bootstrap internal inference on AMD hardware, demonstrating the feasibility of agentic-driven software development on AMD platforms. Microsoft, which dropped AMD after the MI300X due to unreliable Samsung 2023 HBM memory and poor software quality (skipping both MI325X and MI355X), has reversed course and announced a deployment of MI455X Helios. The authors believe OpenAI will be the main customer for these Azure racks. AMD has also announced a disaggregated-inference deal with Cerebras, modeled on the Nvidia-Groq partnership, targeting ultra-fast interactive inferencing.
On the silicon front, AMD has delivered a notable engineering achievement. The MI455X is the first datacenter accelerator to use TSMC's 2nm process on both its compute tiles and the Venice CPU; all competing accelerators remain on 3nm. The package is the largest CoWoS-L module in production at 5.5x reticle size, and AMD is the only adopter of TSMC's SoIC-X hybrid bonding, allowing the company to scale the silicon footprint in the z-dimension as well as x and y. The MI455X achieves a total of 3,470mm² of logic silicon in a single package—"by far the most amount of silicon that is being shipped in a single package." The package integrates 8 compute tiles (XCDs) atop 2 base dies containing SRAM, HBM controllers, and compute fabric, plus two separate I/O dies for off-package communications (UALoE scale-up fabric and inter-chip links).
Financially, AMD is employing creative deal structures to accelerate adoption. AMD has implemented an equity-rebate discount structure with OpenAI and Meta—close to a 105% rebate using financial engineering. The full rebate triggers when AMD stock reaches $600 and once the customers buy enough compute. Combined with Helios's strong cost-per-token economics, the structure results in practically negative unit cost to these customers; as the authors note, "AMD is practically giving away Helios racks and a 5% extra on top of that" to OpenAI (described as "an SF-based nonprofit").
However, the article identifies two critical risks that could undermine this momentum. First, the Helios rack is experiencing a slow production ramp. Unlike Nvidia's Rubin-Oberon architecture, which uses a cableless tray design, Helios relies on a traditional backplane. AMD's weak SerDes design requires up to 85% of the backplane to be retimed, necessitating over 550 Broadcom ethernet retimers per rack. The rack is encountering backplane reliability challenges during the production ramp phase.
Second, and more immediately concerning, AMD faces a severe internal GPU capacity constraint. The chief complaint from AMD engineers is a persistent lack of stable GPU clusters for internal software development and automated testing/continuous integration (CI). This is blocking the pace of progress and preventing the organization from harnessing the potential of AI coding agents, which themselves require GPU resources and testing loops. Specifically, the vLLM team was making progress toward a goal of 90% parity with CUDA on gating tests (the highest-quality tests that prevent buggy PRs from merging) by Advancing AI 2026, until AMD leadership pulled clusters away from the team to address an internal capacity crunch. The planned timeline has slipped to October 2026, with efforts underway to shift back to August/September 2026.
On the Kubernetes Inferencing front—critical since Kubernetes is the layer used by most inference deployments worldwide—the Pollara NIC CI sits at 0% parity with Nvidia's ConnectX Nightly CI. The planned ETA to reach parity by Advancing AI 2026 was missed due to cluster issues. The authors note that the problem is not engineering unwillingness but lack of investment in internal CI capacity. Even with 2,000 additional MI355Xs coming online this month and 6,000 more MI325X/MI355X units later in the year, total capacity will still fall more than an order of magnitude short of the stable long-term clusters Nvidia maintains for internal development. The rise of agentic AI exacerbates the shortage: previously, each human engineer required a couple of GPU nodes for distributed-inference development; now each agent requires GPUs, and each human can run dozens of agents simultaneously, each spawning dozens of sub-agents. Without simulation tools like DynoSim, this capacity gap will widen further. The MI455X compounds the problem by using a completely different ISA (gfx1250) from the MI355 (gfx950), requiring distinct codepaths and kernels that must be tested on both architectures.
The authors credit specific AMD engineers—Hongxia, Chun Fang, HaiShaw, Thomas Wang, Andy Luo, Seungrok, Bill He, Teresa Shan, Parth, Duyi Wang, and Gilbert—for their 10x-level work on ROCm software, noting that many of AMD's best engineers are based in Shanghai and that core pieces of the ROCm stack are built in China. However, they emphasize that the constraint is not engineering talent or will but rather leadership's capacity-planning philosophy. They urge AMD leadership to "re-prioritize" stable cluster allocation and abandon a strategy that shifts clusters between cloud service providers, which forces constant migration and destabilizes development workflows. If AMD can overcome these two risks—production ramp and internal GPU capacity—the authors believe AMD will be well positioned to gain market share despite Nvidia's continued strong growth in an expanding pie.
AMD's competitive position has shifted dramatically over the past year. The company moved from a 0% chance of closing the gap with Nvidia to what the authors now call a "great chance of success," driven by leadership changes, strategic customer wins, and silicon engineering breakthroughs. The securing of Anthropic's 2GW commitment and Microsoft's return after a two-year absence (having dropped AMD after the MI300X due to Samsung HBM reliability and software quality issues) signals growing confidence in AMD's execution. The equity-rebate structure with OpenAI and Meta further sweetens the value proposition, making unit economics favor AMD on a total-cost-of-ownership basis.
However, the article reveals a critical fracture within AMD between its hardware and software momentum. While the MI455X represents a silicon engineering achievement—first to ship 2nm datacenter silicon with the largest on-package silicon footprint (3,470mm²)—the software organization is throttled by internal resource constraints. The lack of stable GPU clusters for development and continuous integration testing has become the binding constraint. vLLM gating tests (the highest-quality tests that prevent buggy code from merging) show 0% parity with Nvidia on Kubernetes inferencing and have missed their Advancing AI 2026 target, now slipping to October 2026. Even with 2,000 MI355Xs and 6,000 MI325X/MI355X units coming online, total internal capacity remains more than an order of magnitude below Nvidia's stable long-term clusters. The rise of agentic AI compounds this: where a human engineer previously needed a couple of GPU nodes, a single agent now requires GPU resources to test code, and agents can spawn dozens of sub-agents—creating geometric demand pressure on finite hardware.
The article positions this as a leadership and philosophy issue, not an engineering capability gap. AMD's software engineers understand the work required; they lack the resources and stable infrastructure to execute at velocity. The authors specifically call out the need for AMD leadership to "re-prioritize" cluster stability and revise capacity-planning strategy, suggesting the constraint is organizational rather than technical.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime