AIToday

Weka cuts AI GPU load with cached tokens, launches storage hardware

VentureBeat AI12h ago
Weka cuts AI GPU load with cached tokens, launches storage hardware

Key takeaway

Weka has released NeuralMesh 6 and its first custom hardware, Wekapod 3, to solve a critical AI production bottleneck: GPU memory is expensive and runs out fast, especially when models must recompute information across long conversations. The new platform uses Weka's Augmented Memory Grid to cache all of an AI model's pre-calculated tokens in much cheaper flash storage, reducing GPU load and potentially letting existing hardware handle more users without additional GPU investment.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Weka released NeuralMesh 6 software and its first custom hardware, Wekapod 3, designed to cache 100% of an AI model's pre-calculated tokens in cheaper flash storage instead of GPU memory, reducing the need to recompute information in long conversations.

  • Why it matters

    GPU memory is the most expensive and fastest-depleting resource in production AI. By offloading token caching to flash storage—orders of magnitude cheaper—Weka's Augmented Memory Grid approach may allow the same hardware to serve more users or handle longer sessions without GPU upgrades.

  • What to watch

    Weka enters a crowded field; Dell, NetApp, Pure Storage, and VAST have all repositioned toward AI infrastructure over the past two years, each claiming purpose-built solutions for this challenge.

In Depth

GPU memory has become the critical constraint in scaling AI inference. Long context windows—the amount of text a model can consider at once—and multi-turn conversations create a compounding problem: models must repeatedly recompute information they have already processed, consuming precious GPU memory and compute capacity. This means fewer concurrent users can be served per GPU, forcing costly hardware expansion. Weka's answer is to bypass this constraint by extending GPU memory into the storage tier. The company's NeuralMesh 6 platform, released alongside Wekapod 3 hardware, implements what Weka calls the Augmented Memory Grid—an architecture that aggregates NAND flash storage to function as GPU memory, but at a fraction of the cost. The key insight is that flash storage can cache 100% of an AI model's pre-calculated tokens, eliminating the need to recompute them on every turn of a conversation. This allows the same GPU to serve more users or handle longer sessions without additional silicon investment. Wekapod 3 marks Weka's entry into custom hardware, a signal that the company believes its software-plus-hardware stack offers an integrated advantage. The timing reflects industry momentum: Dell, NetApp, Pure Storage, and VAST have all repositioned toward AI infrastructure over the past two years, each competing for the same opportunity to reshape how enterprises handle the compute-memory tradeoff. Weka's claim is that it has built from the ground up for this challenge, but the crowded field suggests customers will have multiple architectures to evaluate before settling on a vendor.

Context & Analysis

The bottleneck Weka addresses is real and acute. In production AI systems, GPU memory is not only the costliest resource but also the scarcest—long context windows and multi-turn conversations force models to repeatedly recompute information they have already processed, wasting both GPU memory and compute cycles that could serve additional users or generate new outputs. Weka's strategy is to treat this not as an inevitable constraint but as an opportunity to extend GPU memory with cheaper flash storage technologies, using flash as a high-speed cache layer for pre-calculated tokens. This shift in architecture—from GPU-centric to hybrid GPU-plus-storage—reflects broader industry recognition that traditional infrastructure design leaves compute underutilized when memory becomes the limiting factor. However, Weka does not enter this space alone. Dell, NetApp, Pure Storage, and VAST have all shifted their positioning toward AI infrastructure within the past two years, each claiming similar purpose-built solutions. The emergence of this crowded competitive field suggests both strong market demand and uncertainty about which vendor's approach will dominate, making execution and customer adoption critical for Weka's differentiation.

FAQ

What is Weka's Augmented Memory Grid and how does it work?
It aggregates NAND flash storage to behave like GPU memory at a fraction of the cost, allowing AI models to cache pre-calculated tokens instead of recomputing them during long conversations or multi-turn interactions.
What products did Weka announce?
Weka launched NeuralMesh 6, a software platform, alongside Wekapod 3, the company's first self-designed hardware line.

Get AI news like this every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No discussion yet for this article

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →