AIToday
Large Language ModelsSiliconANGLE AIPublished: Jul 16, 2026, 16:01 JST3 min read

Agentic AI shifts focus from training to storage, memory, power efficiency

Agentic AI shifts focus from training to storage, memory, power efficiency

Key takeaway

  • Agentic inference—longer-running AI agent tasks—is reshaping data center architecture away from a focus on raw compute speed toward storage, memory, and power efficiency as the limiting factors.

  • Companies including Solidigm, AMD, Tensordyne, and d-Matrix presented solutions at RAISE Summit that address bottlenecks in data delivery to GPUs, power consumption, and heterogeneous computing.

  • The shift also highlights capital constraints and data sovereignty as infrastructure concerns, with companies like Argentum AI and Neo4j offering financing and governance solutions alongside hardware.

3 Key Points

  1. What happened

    At RAISE Summit, infrastructure leaders revealed that agentic inference—longer-running AI agent tasks—is creating new bottlenecks in storage, memory, and power consumption. Solidigm's Greg Matson explained that storage has moved into the critical path, becoming "a whole new storage tier that's being created to extend the memory for the system." AMD is optimizing across CPUs, GPUs, and networking rather than individual chips; Tensordyne's Napier inference chip uses logarithmic math to cut power draw to 30 kilowatts (versus 150 kilowatts for a comparable Nvidia system); and d-Matrix is pairing purpose-built accelerators with GPUs for heterogeneous inference in production.

  2. Why it matters

    As enterprises move from single-task AI workloads to long-running agentic systems, the infrastructure bottleneck has shifted from raw compute speed to keeping GPUs continuously fed with data and managing power costs. Storage positioned near accelerators is now critical to prevent idle GPU time—which wastes capital since GPUs are the most expensive part of the infrastructure. This reshaping of the AI stack means businesses cannot rely on GPU capacity alone; they must rethink storage, memory staging, and power architecture to stay competitive.

  3. What to watch

    Data sovereignty and financing are emerging as infrastructure concerns alongside hardware. Argentum AI is addressing capital constraints by securing customer contracts first, then financing construction—a model that treats "power, compute and capital" as an integrated product. Agentcy Labs and Neo4j are advancing knowledge graphs as a way to give enterprises deterministic control and explainability alongside large language models, positioning sovereignty (territorial, operational, and stack control) as a core infrastructure requirement for agentic deployments.

Ask the AI about this article →

Context & Analysis

The article documents a fundamental shift in how enterprises design AI infrastructure. The race to scale training—which dominated the past two years—has given way to a phase where inference, particularly agentic inference (longer-running AI agent tasks), is reshaping priorities. Agentic workloads are exposing new bottlenecks that were previously secondary: storage capacity and bandwidth, power efficiency, and the need for heterogeneous computing (specialized accelerators working alongside GPUs rather than replacing them).

This shift is not merely a technical optimization but a strategic reorientation. As Greg Matson explained, hyperscalers are now treating storage as an active extension of GPU memory, not an isolated component. The imperative is to keep GPUs "humming 100% of the time generating tokens," because idle GPU time wastes the largest capital investment. This logic has also accelerated the adoption of power-efficient architectures—Tensordyne's use of logarithmic math to replace multiplications with additions is a concrete example of rethinking silicon design for the constraints of continuous inference rather than batch training.

Beyond hardware, the article reveals that capital and governance have become infrastructure concerns. Argentum AI's financing-first model and the emphasis on data sovereignty through knowledge graphs indicate that enterprises building agentic systems are no longer purchasing isolated components but assembling integrated stacks that account for speed of deployment, cost of capital, and control over proprietary data. This reframing—treating capital, sovereignty, and architecture as co-equal challenges—signals a maturation of the agentic AI market from proof-of-concept to production scale.

FAQ

How much power does Tensordyne's Napier chip use compared to Nvidia?
A 72-chip Napier pod draws 30 kilowatts, compared with 150 kilowatts for a comparable Nvidia Corp. system, according to Tensordyne co-founder Gilles Backhus.
What problem is Argentum AI solving in AI infrastructure?
Argentum AI addresses capital constraints by securing customers and contracted revenue before committing capital to construction, treating financing as a core part of the deployment stack alongside power and compute.
Why is storage becoming critical for agentic inference?
As agentic inference expands from individual prompts to longer-running sessions, the volume of context data can exceed GPU memory capacity, making high-capacity storage positioned near accelerators essential to keep GPUs continuously generating tokens and to prevent wasting expensive GPU capacity waiting for data.
SiliconANGLE AIRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Visko raises $10M, launches live AI video model OrbisSiliconANGLE AI · 2h ago
  • Runway unveils Solaris, an AI that generates app interfaces in real timeTHE DECODER · 2h ago
  • Google AI Search flags Facebook users as dangerTHE DECODER · 2h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleXi to keynote Shanghai AI conference July 17