AIToday

Tensormesh and AMD let fewer GPUs run more AI models

Yahoo Finance AI6h ago
Tensormesh and AMD let fewer GPUs run more AI models

Key takeaway

Tensormesh and AMD have jointly announced that their combined KV cache and virtual memory technologies allow enterprises to serve more AI models on fewer GPUs without sacrificing speed or throughput. In testing, the solution cut response time by nearly 7× on large document sets, maintained steady output at ~48 tokens per second under growing load, and doubled model density on the same hardware—all while reducing costs and avoiding the need to purchase GPUs with larger, more expensive memory.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Tensormesh and AMD announced a collaboration in which Tensormesh's LMCache (a KV cache solution) works with AMD's virtual memory offering to let enterprises run more AI models on the same number of GPUs while keeping inference fast. Testing on Dell servers with 8x AMD MI355 GPUs showed time-to-first-token dropped from 3.4 seconds to under half a second on a 300 GB document set, output throughput stayed steady at ~48 tokens per second even as workload grew, and model density doubled on the same hardware.

  • Why it matters

    Running more models on fewer GPUs cuts infrastructure costs for enterprises processing large document sets, and avoids the current memory supply crisis by eliminating the need to buy GPUs with larger memory. Instead of building new hardware, the solution reuses existing infrastructure by offloading KV cache (a form of temporary memory AI uses during inference) onto cheaper storage like SSDs and remote systems.

  • What to watch

    The collaboration is built on AMD's recent strategic investment in Tensormesh; AMD also announced a new breakthrough in virtual GPU memory management that powers the integration.

In Depth

On July 23, 2026, Tensormesh and AMD announced a partnership to optimize AI inference by allowing enterprises to run more models on fewer GPUs. The solution combines Tensormesh's LMCache (a KV cache management system) with AMD's virtual memory offering and GPU technology, specifically AMD Live Context Virtualization components. LMCache enables KV cache—temporary memory that AI models use during inference—to overflow from the GPU's high-bandwidth memory (HBM) onto a multi-tier storage hierarchy that includes DRAM, SSDs, and remote storage.

Testing was conducted on Dell servers equipped with 8x AMD MI355 GPUs and Dell storage. The results were substantial. On a 300 GB document set using the Kimi-K2.6 model, time-to-first-token (the time before the model produces its first response token) dropped from 3.4 seconds to under half a second—a nearly 7× improvement—by reusing cached KV data already stored in DRAM and NFS instead of recomputing it from scratch. Output throughput remained stable at approximately 48 tokens per second regardless of workload size, while unoptimized inference systems experienced a nearly 40% drop in throughput under the same conditions. Hardware model density doubled on the same infrastructure.

The economic case is clear: enterprises can reuse existing GPU infrastructure without buying new GPUs with larger memory, directly lowering costs and sidestepping the industry's current memory supply crisis. Customers can also maximize efficiency by reusing KV cache chunks stored for short-term memory virtualization later for prefix or non-prefix KV cache matching. Tensormesh CEO and LMCache co-creator Junchen Jiang stated: "Together, we're greatly enhancing the memory that inference engines can access for model weights and KV cache, using all of the memory resources on each node." AMD's Vice President of AI Software Anush Elangovan added: "We recognize Tensormesh and LMCache as KV cache management leaders. And we're thrilled to announce AMD's breakthrough in virtual GPU memory management, amplified by LMCache and Tensormesh." The collaboration builds on AMD's recent strategic investment in Tensormesh.

Context & Analysis

The collaboration addresses a structural problem in enterprise AI infrastructure: GPU memory (high-bandwidth memory, or HBM) is expensive and in short supply, yet inference workloads often need fast access to large amounts of temporary data (KV cache) generated during model computation. Rather than forcing customers to buy GPUs with more memory built in—the traditional approach that has contributed to the current memory supply crisis—Tensormesh and AMD's solution allows KV cache to overflow into cheaper, more abundant storage tiers: DRAM, SSDs, and remote storage. The key innovation is LMCache, which coordinates this multi-tier cache intelligently, allowing the system to reuse cached data for future requests (both prefix and non-prefix KV cache matching) rather than recomputing it from scratch.

The test results validate the economic trade-off: a nearly 7× speedup in time-to-first-token on large document sets, combined with doubled model density on the same hardware, means enterprises can consolidate workloads and reduce capital expenditure. The fact that output throughput remained constant even as load grew—while unoptimized systems degraded by 40%—suggests the solution scales well for real-world oversubscribed environments. AMD's recent strategic investment in Tensormesh and its simultaneous announcement of a virtual GPU memory management breakthrough indicate that this is not an isolated experiment but part of a broader AMD strategy to compete in the inference optimization market.

FAQ

How much faster did time-to-first-token improve in testing?
On a 300 GB document set using Kimi-K2.6, time-to-first-token dropped from 3.4 seconds to under half a second, a nearly 7× improvement.
What hardware was used to test the solution?
Testing was performed using Dell servers with 8x AMD MI355 GPUs and Dell storage.
How did output throughput perform under increasing load?
Output throughput held steady at ~48 tokens per second regardless of workload size, while unoptimized inference throughput fell by nearly 40% under the same load.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →