
TensorWave has integrated AMD's Helios rackscale solution into its AI cloud platform, combining 72 high-end GPUs with memory and networking technology to deliver exaflop-class performance for large-scale AI training and inference. The solution addresses growing demand for high-performance, memory-intensive workloads and offers customers a path to scale infrastructure efficiently from rack to cluster.
Summaries like this, in your inbox every morning.
Sign up free →What happened
TensorWave, an all-AMD AI cloud provider, announced it has adopted the AMD Helios rackscale solution, which integrates 72 AMD Instinct MI455X GPUs, 6th Gen AMD EPYC CPUs, AMD Pensando networking, and AMD ROCm software into a single unified platform for large-scale inference and frontier-model training workloads.
Why it matters
The solution delivers exaflop-class AI performance—up to 2.9 exaFLOPS peak 4-bit and 1.4 exaFLOPS peak 8-bit at rack scale, with 31 TB of HBM4 memory and 1.7 PB/s of memory bandwidth—addressing the accelerating demand for large-scale inference and frontier-model development that most cloud providers struggle to support efficiently.
What to watch
The platform enables customers to scale from a single rack to multi-rack clusters with consistent deployment. Featherless AI, one of the first adopters cited in the announcement, emphasized confidence in TensorWave's deep AMD infrastructure expertise and ability to support future AMD hardware releases.
On July 23, 2026, TensorWave announced its adoption of AMD's Helios rackscale solution to power its AI cloud platform. The solution combines 72 AMD Instinct MI455X GPUs with 6th Gen AMD EPYC CPUs, AMD Pensando networking technology, and AMD ROCm software into a unified platform designed for large-scale AI inference, frontier-model training, and fine-tuning.
The hardware delivers substantial performance: at rack scale, the solution achieves up to 2.9 exaFLOPS peak 4-bit (OCP MXFP4) and 1.4 exaFLOPS peak 8-bit (OCP MXFP8), with 31 TB of HBM4 memory and 1.7 PB/s of memory bandwidth. Each individual MI455X GPU contributes up to 40 PFLOPs peak 4-bit, 20 PFLOPs peak 8-bit, 432 GB HBM4, and 23.3 TB/s memory bandwidth. The platform comprises 18 ORW-aligned 4-GPU compute trays (totaling 72 GPUs) that form a unified architecture capable of being virtualized for consistent deployment and expansion across clusters. The AMD Pensando Vulcano 800 AI NIC provides the high-bandwidth, low-latency connectivity needed to support multi-trillion-parameter training.
Andrew Dieckmann, Corporate Vice President and General Manager of AMD's Data Center GPU Business Unit, emphasized that AMD Helios delivers "leadership performance, performance per watt, and cost per token for the next wave of AI," positioning it as a solution for customers facing accelerating demand for large-scale inference and frontier-model development. TensorWave's Chief Growth Officer and Co-Founder Jeff Tatarchuk stated that "AMD Helios brings compute, memory, and networking together at rack scale, enabling our customers to build AI faster, more efficiently, and with better economics at scale."
Featherless AI, one of the early adopters highlighted in the announcement, noted that it selected TensorWave specifically because of the company's deep expertise with AMD infrastructure. CEO Eugene Cheah said, "That expertise and relationship with AMD gives us confidence not just in what we're running today, but in knowing they'll be ready for whatever AMD ships next," underscoring the value of vendor specialization in an environment of rapid hardware innovation.
TensorWave's adoption of AMD Helios reflects a broader industry shift toward purpose-built infrastructure for large-scale AI workloads. The solution unifies compute (GPUs and CPUs), memory (31 TB HBM4), and networking (Pensando) into a single rack-scale building block—a design approach that avoids the fragmentation and integration challenges that plague multi-vendor cloud setups. By standardizing on AMD's entire stack, TensorWave positions itself to compete directly with hyperscalers that may have broader hardware options but less optimized coordination across components.
The performance metrics underscore the hardware's orientation toward frontier-scale work: 2.9 exaFLOPS at rack scale with 1.7 PB/s memory bandwidth is substantial enough to train multi-trillion-parameter models and run high-throughput inference. The emphasis on "exaflop-class" performance and the inclusion of the Pensando Vulcano 800 AI NIC for low-latency, high-bandwidth connectivity suggests TensorWave is targeting customers who cannot tolerate the latency and efficiency losses of general-purpose cloud infrastructure. Early adopter Featherless AI's rationale—confidence in TensorWave's AMD expertise and readiness for future releases—implies that deep specialization in a single vendor's ecosystem may be a competitive advantage when that vendor innovates rapidly.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion



Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime
1 minute a day. The AI essentials.
200+ sources · Email / LINE / Slack