
Liqid and AMD have jointly unveiled a new AI infrastructure platform that pools up to 30 AMD Instinct MI350P GPUs in a single server, delivering 69 PFLOPS of AI inference performance. The collaboration aims to reduce the cost and complexity of deploying large language models by offering 65% lower deployment costs and 50% lower power consumption compared to traditional setups, without requiring datacenter redesigns or specialized cooling.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Liqid announced a strategic collaboration with AMD to deliver AI infrastructure combining up to 30x AMD Instinct MI350P GPUs with Liqid's GPU pooling platform. The resulting system, called Liqid UltraStack 30, delivers 69 PFLOPS (FP8) performance, 4.3 TB of GPU HBM memory, and consumes approximately 22 kW of power in a single server.
Why it matters
The platform lets enterprise and cloud customers scale GPU resources without rebuilding their datacenters or installing specialized cooling. By pooling GPUs through software, it targets up to 65% lower deployment cost and 50% lower power consumption for large AI models, while enabling multiple models to run in parallel on a single system.
What to watch
The system supports native Kubernetes for orchestration and is available now; Liqid positions it for enterprise AI inference workloads where token economics (tokens per second, per dollar, and per watt) determine feasibility.
On July 23, 2026, Liqid announced a strategic collaboration with AMD centered on delivering next-generation AI infrastructure built around AMD's Instinct MI350P GPUs. Liqid, which specializes in software-defined memory and GPU pooling infrastructure, has combined its pooling solutions with AMD's PCIe-based Instinct GPUs to create the Liqid UltraStack 30, a system that can accommodate up to 30 MI350P GPUs in a single server.
The UltraStack 30 is configured with a dual-socket AMD EPYC 9005 Series CPU, 30x AMD Instinct MI350P GPUs, and delivers 69 PFLOPS (FP8) of AI inference performance. The system scales 4.3 TB of aggregate HBM3E memory—enabling the deployment of very large AI models—and consumes approximately 22 kW of total power. Critically, the platform does not require specialized cooling or a complete datacenter redesign, making it accessible to enterprise and cloud customers operating within existing infrastructure constraints.
Liqid CEO Rick Hegberg stated that "the AMD Instinct MI350P Series gives us the ideal PCIe-based GPU to build the next generation of AI infrastructure, allowing us to deliver solutions tuned for enterprise AI inference where utilization and cost per token decide the economics." Suresh Andani, corporate vice president of Compute and Enterprise AI Group at AMD, echoed this sentiment, emphasizing that the MI350P's design allows "customers to scale GPU resources on demand to achieve higher utilization, lower infrastructure costs, and industry-leading AI inference economics."
The solution targets significant operational improvements: up to 65% lower deployment cost and 50% lower power consumption for frontier and large-context models. The platform supports multiple parallel model deployments through native Kubernetes integration, allowing organizations to maximize GPU utilization across mixed workloads. By pooling GPUs through software rather than relying on custom interconnects or specialized infrastructure, Liqid and AMD position the UltraStack 30 as a practical path for enterprises seeking to deploy large AI models without complete infrastructure overhauls.
Liqid's GPU pooling technology addresses a core challenge in modern AI infrastructure: as inference workloads grow in scale and complexity, organizations struggle to balance performance, utilization, and cost without overhauling their datacenters. The partnership with AMD leverages the MI350P's PCIe-based design—which does not require the specialized interconnects or cooling infrastructure that other high-end accelerators demand—to make large-scale GPU pooling practical for enterprises. By enabling a single server to dynamically pool and scale 30 GPUs through software, the UltraStack 30 eliminates the need for complete datacenter redesigns while preserving the ability to deploy frontier and large-context models efficiently.
The stated economics—targeting 65% lower deployment cost and 50% lower power consumption—represent a significant competitive claim in a market where token cost is increasingly the limiting factor for AI service providers. The inclusion of native Kubernetes support and the ability to run multiple models in parallel suggest the platform is designed for production cloud and enterprise environments where multi-tenant and mixed-workload scenarios are common. Liqid's positioning as "the leader in GPU pooling and scaling" and AMD's endorsement of the pairing indicate both vendors believe PCIe-based pooling, rather than custom interconnects, is the path to scalable, cost-effective AI inference infrastructure.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion


Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime
1 minute a day. The AI essentials.
200+ sources · Email / LINE / Slack