AIToday

AMD launches Helios AI rack with 30% better token efficiency

Yahoo Finance AI6h ago
AMD launches Helios AI rack with 30% better token efficiency

Key takeaway

AMD launched AMD Helios, a new rack-scale AI solution that delivers up to 30% more inference tokens per dollar than competing systems, now in production with leading AI companies including OpenAI, Anthropic, and Meta. The solution combines 72 high-performance GPUs and 18 CPUs with AMD's Pensando networking and ROCm software to address growing demand for inference and agentic AI workloads at scale.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    AMD unveiled its full AI infrastructure portfolio at Advancing AI 2026, including 6th Gen AMD EPYC CPUs, AMD Instinct MI400 Series GPUs, and AMD Helios rackscale solutions now in production. AMD Helios is built with 72 high-performance AMD Instinct MI455X GPUs and 18 6th Gen AMD EPYC "Venice" CPUs connected by AMD Pensando networking.

  • Why it matters

    AMD Helios delivers up to 30% more inference tokens per dollar than competing solutions, directly improving cost efficiency for AI companies scaling inference and agentic workloads. Leading AI labs including OpenAI, Anthropic, Meta, Microsoft, and Oracle are already choosing AMD Helios, indicating customer confidence in the open, full-stack architecture.

  • What to watch

    AMD Helios systems are now in production and will be available from major OEMs including Bull, HPE, Lenovo, and Supermicro, as well as infrastructure partners Sanmina and Wiwynn. AMD estimates its total addressable market will reach ~$2 trillion(約320兆円) in 2030 as AI demand spans data center, PCs, edge, and embedded processors.

In Depth

On July 23, 2026, AMD announced a comprehensive AI infrastructure refresh at Advancing AI 2026, with AMD Helios as its flagship product. AMD Helios is a rack-scale AI solution designed from the ground up for large-scale AI inference and agentic workloads. The system integrates 72 AMD Instinct MI455X GPUs (the latest in AMD's MI400 series) and 18 6th Gen AMD EPYC "Venice" CPUs, connected by AMD's proprietary Pensando networking (handling front-end, scale-up, and scale-out traffic) and optimized by AMD's ROCm open software stack.

The key performance claim is that AMD Helios delivers up to 30% more inference tokens per dollar than the leading competitive solution, a metric that directly addresses the operational cost of running large-scale inference workloads. For AI companies deploying at gigawatt scale, a 30% efficiency gain translates to substantial capital and operating expense savings.

AMD has secured commitments from major AI infrastructure consumers. OpenAI, Anthropic, Meta, Microsoft, Oracle, and several specialized infrastructure providers (HUMAIN, Tensorwave, Vultr, Cirrascale) are publicly backing AMD Helios. Hardware will be sold through established OEM partners including Bull, HPE, Lenovo, and Supermicro, plus infrastructure specialists Sanmina and Wiwynn, enabling broad distribution. AMD also unveiled complementary products including 6th Gen AMD EPYC CPUs, AMD Ryzen AI Embedded X100 processors, and the AMD Kria AI System-on-Module and Robotics Developer Platform, signaling a full-stack play across data center, PC, edge, and embedded markets.

Dr. Lisa Su, AMD's chair and CEO, framed the announcement around a broader industry shift: as AI moves from training frontier models to inference and agentic AI, customer demand spans different compute requirements across the stack. AMD's bet is that an open, modular platform—rather than a vertically integrated one—will attract customers seeking flexibility and choice. AMD projects its total addressable market will grow to ~$2 trillion(約320兆円) by 2030, driven by accelerating compute demand across data center, PCs, edge, and embedded processors.

Context & Analysis

AMD's announcement at Advancing AI 2026 reflects a strategic shift in the AI market from training-focused infrastructure to inference and agentic workloads. By unveiling a fully integrated rack architecture—AMD Helios—the company is positioning itself directly against specialized AI infrastructure vendors. The 30% token-per-dollar advantage is a concrete performance metric designed to appeal to cost-conscious cloud providers and AI labs deploying large inference clusters.

The timing and roster of partners underscore AMD's ambition to establish an open alternative to Nvidia's dominant position in AI accelerators. OpenAI, Anthropic, Meta, and Microsoft are among the world's largest AI infrastructure spenders; their public endorsement of AMD Helios signals willingness to diversify GPU suppliers. AMD's emphasis on an "open, full-stack" platform—including custom networking (Pensando), software (ROCm), and a choice of OEM partners—is a deliberate contrast to Nvidia's vertically integrated ecosystem.

AMD's stated TAM projection of ~$2 trillion(約320兆円) in 2030 across data center, PC, edge, and embedded processors reflects a broader bet that compute demand for AI will span multiple form factors and deployment scales, not just cloud data centers. This framing positions AMD's diverse chip portfolio as a long-term advantage if inference and agentic AI workloads proliferate across enterprise, edge, and embedded systems.

FAQ

What is AMD Helios and what hardware does it include?
AMD Helios is a rackscale AI solution built with 72 high-performance AMD Instinct MI455X GPUs and 18 6th Gen AMD EPYC "Venice" CPUs, connected by AMD Pensando front-end, scale-up and scale-out networking, and accelerated by AMD ROCm open software.
Who is already using AMD Helios?
Leading AI labs and cloud providers including OpenAI, Anthropic, Meta, Microsoft, Oracle, HUMAIN, Tensorwave, Vultr, and Cirrascale are choosing AMD Helios. Systems will be available from OEMs including Bull, HPE, Lenovo and Supermicro, as well as infrastructure partners Sanmina and Wiwynn.
How does AMD Helios compare to competing AI racks?
AMD Helios delivers up to 30% more inference tokens per dollar than the leading competitive solution, maximizing output from every rack deployed.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →