
Agentic inference—longer-running AI agent tasks—is reshaping data center architecture away from a focus on raw compute speed toward storage, memory, and power efficiency as the limiting factors.
Companies including Solidigm, AMD, Tensordyne, and d-Matrix presented solutions at RAISE Summit that address bottlenecks in data delivery to GPUs, power consumption, and heterogeneous computing.
The shift also highlights capital constraints and data sovereignty as infrastructure concerns, with companies like Argentum AI and Neo4j offering financing and governance solutions alongside hardware.
What happened
At RAISE Summit, infrastructure leaders revealed that agentic inference—longer-running AI agent tasks—is creating new bottlenecks in storage, memory, and power consumption. Solidigm's Greg Matson explained that storage has moved into the critical path, becoming "a whole new storage tier that's being created to extend the memory for the system." AMD is optimizing across CPUs, GPUs, and networking rather than individual chips; Tensordyne's Napier inference chip uses logarithmic math to cut power draw to 30 kilowatts (versus 150 kilowatts for a comparable Nvidia system); and d-Matrix is pairing purpose-built accelerators with GPUs for heterogeneous inference in production.
Why it matters
As enterprises move from single-task AI workloads to long-running agentic systems, the infrastructure bottleneck has shifted from raw compute speed to keeping GPUs continuously fed with data and managing power costs. Storage positioned near accelerators is now critical to prevent idle GPU time—which wastes capital since GPUs are the most expensive part of the infrastructure. This reshaping of the AI stack means businesses cannot rely on GPU capacity alone; they must rethink storage, memory staging, and power architecture to stay competitive.
What to watch
Data sovereignty and financing are emerging as infrastructure concerns alongside hardware. Argentum AI is addressing capital constraints by securing customer contracts first, then financing construction—a model that treats "power, compute and capital" as an integrated product. Agentcy Labs and Neo4j are advancing knowledge graphs as a way to give enterprises deterministic control and explainability alongside large language models, positioning sovereignty (territorial, operational, and stack control) as a core infrastructure requirement for agentic deployments.
Ask the AI about this article →
The article documents a fundamental shift in how enterprises design AI infrastructure. The race to scale training—which dominated the past two years—has given way to a phase where inference, particularly agentic inference (longer-running AI agent tasks), is reshaping priorities. Agentic workloads are exposing new bottlenecks that were previously secondary: storage capacity and bandwidth, power efficiency, and the need for heterogeneous computing (specialized accelerators working alongside GPUs rather than replacing them).
This shift is not merely a technical optimization but a strategic reorientation. As Greg Matson explained, hyperscalers are now treating storage as an active extension of GPU memory, not an isolated component. The imperative is to keep GPUs "humming 100% of the time generating tokens," because idle GPU time wastes the largest capital investment. This logic has also accelerated the adoption of power-efficient architectures—Tensordyne's use of logarithmic math to replace multiplications with additions is a concrete example of rethinking silicon design for the constraints of continuous inference rather than batch training.
Beyond hardware, the article reveals that capital and governance have become infrastructure concerns. Argentum AI's financing-first model and the emphasis on data sovereignty through knowledge graphs indicate that enterprises building agentic systems are no longer purchasing isolated components but assembling integrated stacks that account for speed of deployment, cost of capital, and control over proprietary data. This reframing—treating capital, sovereignty, and architecture as co-equal challenges—signals a maturation of the agentic AI market from proof-of-concept to production scale.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Visko raised $10 million in pre-seed funding from Llama Ventures and opened public access to its first foundat…
AI company Runway has unveiled Solaris, the first model in a new category it calls "Interface World Models." I…

Google's AI search gave advice to call emergency services for users alone with an African, Indian, or Pakistan…

John Deere introduced JD, a conversational AI tool that lets farmers ask open-ended questions about their hist…

Nvidia CEO Jensen Huang said on Fox Business that AI is creating 'hundreds of thousands' of jobs, including in…

Israeli startup DataAgent Ltd