AIToday
Large Language ModelsAI Business & IndustrySiliconANGLE AIPublished: Oct 7, 2026, 04:00 JST

Vast: Alon Horev says AI agent memory moving to tiered storage

Vast: Alon Horev says AI agent memory moving to tiered storage

3 Key Points

  1. What happened

    Alon Horev, co-founder and CTO of Vast, said at Fully Connected 2026 that agent sessions of half a million tokens can use one-tenth to one-twentieth of GPU memory, so Vast tiers KV cache across GPU memory, CPU memory, and persistent media holding petabytes, orchestrated by Nvidia Corp.'s Dynamo.

  2. Why it matters

    Long agent sessions can outlast GPU memory, so offloading KV cache to storage lets enterprises avoid repeat recalculation and schedule across a fleet, Horev said. He also argued agent conversations are governance and training assets.

  3. What to watch

    The approach hinges on whether enterprises with thousands of agents handling sensitive data actually adopt tiered storage and record everything their agents do. Vast has launched a confidential computing service for sensitive workloads.

WHO IT HITSEnterprise infrastructure teams running AI agents in production and storage buyers evaluating how to hold persistent agent context will face decisions about GPU memory limits and KV cache offloading. IT and compliance staff, especially those handling sensitive data, may need to record and retain agent actions.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Vast's argument rests on a shift in how AI inference is expected to work. Horev described sessions that hold a KV cache in GPU memory and then go idle when the agent compiles code, tests software, or pauses for a person. Offloading a session to storage keeps that memory wall from forcing a repeat calculation, he said. The tiering order he outlined runs from GPU memory to CPU memory on the same machine to persistent media holding petabytes of KV cache, with Nvidia Corp.'s Dynamo software orchestrating the process.

Horev framed this as treating inference as a distributed problem across a fleet of machines. Moving a session from a busy GPU to a less busy one, and reading KV cache either over the network or from Vast, gives operators more optionality and more optimized scheduling, he said. That matters differently from ordinary context handling, he argued, because enterprise agents also need shared knowledge that persists across sessions, including long-term memory of past conversations.

He also tied agent memory to governance and training. Companies running thousands of agents that handle sensitive data and act for customers need to record everything those agents do and retain it for a set period, he said, and Vast has launched a confidential computing service for sensitive workloads. The stakes appear to hinge on whether enterprises adopt this tiered model at fleet scale and whether persistent KV cache proves to be as valuable as Horev suggests. Horev's claim that agent conversations are "gold" for fine-tuning, training, or purpose-built models is his view, not a stated benchmark.

FAQ
Why is GPU memory a problem for AI agents?
Horev said each long-running session holds its KV cache in GPU memory, and a session of half a million tokens can take one-tenth to one-twentieth of a GPU's memory. Offloading those sessions to storage lets the system avoid repeat recalculation.
What does Vast's tiered memory approach actually use?
It uses GPU memory first, then CPU memory on the same machine, then persistent media that can hold petabytes of KV cache. Nvidia Corp.'s Dynamo software orchestrates the process.
What else did Horev say about agent data?
Horev said agent conversations are also useful for fine-tuning, training, or creating purpose-built models, and that enterprises need to record everything their agents do and retain it for a set period. Vast has also launched a confidential computing service for sensitive workloads.
SiliconANGLE AIRead Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleTaiwan's NSTC seeks power tech for AI data centers