AIToday

Dell, AMD push modular AI infrastructure for enterprise scaling

Top Companies AI — US (1/2)4h agoSend on LINE
Dell, AMD push modular AI infrastructure for enterprise scaling

Key takeaway

Dell and AMD have jointly developed a modular AI platform designed to help enterprises transition AI projects from proof of concept to full production deployment. The platform addresses three key enterprise pain points—runaway token-based pricing costs, data governance, and security at scale—by letting organizations start with small, pre-validated modular units and scale incrementally on the same infrastructure without rearchitecting. This approach reflects a broader shift in enterprise AI adoption toward on-premises deployments and flexible deployment models that reduce dependency on cloud-based APIs.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Dell and Advanced Micro Devices Inc. designed the Dell AI Platform with AMD to help enterprises scale AI deployments from proof of concept to production using modular, composable units of storage, compute, networking, and AMD Instinct and EPYC CPUs with ROCm software.

  • Why it matters

    Enterprises moving AI to production face three critical challenges—uncontrolled token-based pricing costs, data governance and RAG pipeline setup, and security concerns at scale. Modular infrastructure lets organizations start small and expand on the same platform without redesigning their systems, avoiding costly rearchitecture.

  • What to watch

    Dell's modular AI Factory approach allows customers to test and validate integrated components before scaling; the framework supports both on-premises deployments (so enterprises become their own token generators) and hybrid models tailored to different workloads.

In Depth

Enterprise adoption of generative AI has reached an inflection point. While proof-of-concept projects often succeed with small user groups and limited scope, moving AI into production at enterprise scale introduces operational and financial complexities that many organizations were not prepared to face.

Varun Chhabra, senior vice president of product marketing and infrastructure solutions group at Dell Technologies, highlighted three interconnected challenges in an exclusive interview at the AMD Advancing AI event. First, the cost of native token-based pricing—the cloud model in which customers pay per API call or token processed—becomes "really, really out of control" as high-value users increase their usage. Second, organizations must establish data governance frameworks and implement technologies such as RAG (retrieval-augmented generation) pipelines to ensure models receive appropriate data while maintaining control and traceability. Third, security and governance become exponentially more complex at scale, as enterprises must guard against unintended consequences in deployed AI systems.

In response, Dell and AMD jointly designed the Dell AI Platform to enable enterprises to start with small, modular deployments and scale incrementally without rebuilding their infrastructure. The platform's modular AI Factory approach provides pre-integrated, pre-validated composable units encompassing storage, compute, networking, AMD Instinct accelerators, and EPYC CPUs running ROCm software. According to Chhabra, customers can "start small with these composable units of storage, compute, networking, pre-integrated with AMD Instinct and EPYC CPUs with the ROCm software," test and validate the configuration, and then "scale in a modular way with the same frameworks" as they see value. This stands in contrast to the traditional cloud-API model, where growing inference workloads force customers into public cloud infrastructure or require costly rearchitecture of on-premises systems.

The timing reflects a broader industry trend: as inference workloads consume an increasing share of AI compute resources, enterprises are exploring on-premises deployments to reduce dependency on cloud-based APIs and become their own token generators. Rather than adopting a single monolithic approach, large enterprises are expected to deploy different AI infrastructure and models tailored to different workloads—a strategy for which modular, validated building blocks are a prerequisite.

Context & Analysis

The shift from enterprise AI pilots to production workloads has exposed a fundamental mismatch between cloud-native consumption models and large-scale operational needs. Token-based pricing, designed for API-first architectures, becomes prohibitively expensive when applied to inference-heavy production workloads serving thousands of users simultaneously. This cost pressure, combined with data sovereignty and governance requirements that many enterprises cannot meet in public cloud environments, has created an opening for modular, on-premises alternatives.

Dell and AMD's approach reflects a market-wide recognition that no single deployment model fits all enterprise AI use cases. By designing the Dell AI Platform as a composable set of validated building blocks—compute, storage, networking, CPU, and software—rather than a monolithic appliance, the vendors enable enterprises to match infrastructure to specific workload requirements. This modularity also reduces the operational risk of migration: organizations can prove value with a small deployment and expand with confidence that the tested foundation will scale, avoiding the costly rearchitecture that has plagued enterprise AI rollouts to date.

FAQ

What are the main challenges enterprises face when scaling AI from proof of concept to production?
According to Varun Chhabra, senior vice president of product marketing and infrastructure solutions group at Dell Technologies, the three key challenges are uncontrolled costs of native token-based pricing, data governance and ensuring AI models receive the right data through technologies such as RAG pipelines, and security and governance concerns around unintended consequences when scaling.
How does the Dell AI Platform with AMD help enterprises scale?
The platform uses modular, composable units of storage, compute, networking, AMD Instinct and EPYC CPUs, and ROCm software that are pre-integrated and tested. Customers can start small and scale on the same platform without rearchitecting their infrastructure, and the modular AI Factory approach lets them validate components before expanding.
Why are enterprises exploring on-premises AI deployments?
As inference workloads grow and consume an increasing share of AI compute, enterprises are exploring on-premises deployments to become their own token generators instead of depending on cloud-based APIs.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime