
Weka has released NeuralMesh 6 and its first custom hardware, Wekapod 3, to solve a critical AI production bottleneck: GPU memory is expensive and runs out fast, especially when models must recompute information across long conversations. The new platform uses Weka's Augmented Memory Grid to cache all of an AI model's pre-calculated tokens in much cheaper flash storage, reducing GPU load and potentially letting existing hardware handle more users without additional GPU investment.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Weka released NeuralMesh 6 software and its first custom hardware, Wekapod 3, designed to cache 100% of an AI model's pre-calculated tokens in cheaper flash storage instead of GPU memory, reducing the need to recompute information in long conversations.
Why it matters
GPU memory is the most expensive and fastest-depleting resource in production AI. By offloading token caching to flash storage—orders of magnitude cheaper—Weka's Augmented Memory Grid approach may allow the same hardware to serve more users or handle longer sessions without GPU upgrades.
What to watch
Weka enters a crowded field; Dell, NetApp, Pure Storage, and VAST have all repositioned toward AI infrastructure over the past two years, each claiming purpose-built solutions for this challenge.
GPU memory has become the critical constraint in scaling AI inference. Long context windows—the amount of text a model can consider at once—and multi-turn conversations create a compounding problem: models must repeatedly recompute information they have already processed, consuming precious GPU memory and compute capacity. This means fewer concurrent users can be served per GPU, forcing costly hardware expansion. Weka's answer is to bypass this constraint by extending GPU memory into the storage tier. The company's NeuralMesh 6 platform, released alongside Wekapod 3 hardware, implements what Weka calls the Augmented Memory Grid—an architecture that aggregates NAND flash storage to function as GPU memory, but at a fraction of the cost. The key insight is that flash storage can cache 100% of an AI model's pre-calculated tokens, eliminating the need to recompute them on every turn of a conversation. This allows the same GPU to serve more users or handle longer sessions without additional silicon investment. Wekapod 3 marks Weka's entry into custom hardware, a signal that the company believes its software-plus-hardware stack offers an integrated advantage. The timing reflects industry momentum: Dell, NetApp, Pure Storage, and VAST have all repositioned toward AI infrastructure over the past two years, each competing for the same opportunity to reshape how enterprises handle the compute-memory tradeoff. Weka's claim is that it has built from the ground up for this challenge, but the crowded field suggests customers will have multiple architectures to evaluate before settling on a vendor.
The bottleneck Weka addresses is real and acute. In production AI systems, GPU memory is not only the costliest resource but also the scarcest—long context windows and multi-turn conversations force models to repeatedly recompute information they have already processed, wasting both GPU memory and compute cycles that could serve additional users or generate new outputs. Weka's strategy is to treat this not as an inevitable constraint but as an opportunity to extend GPU memory with cheaper flash storage technologies, using flash as a high-speed cache layer for pre-calculated tokens. This shift in architecture—from GPU-centric to hybrid GPU-plus-storage—reflects broader industry recognition that traditional infrastructure design leaves compute underutilized when memory becomes the limiting factor. However, Weka does not enter this space alone. Dell, NetApp, Pure Storage, and VAST have all shifted their positioning toward AI infrastructure within the past two years, each claiming similar purpose-built solutions. The emergence of this crowded competitive field suggests both strong market demand and uncertainty about which vendor's approach will dominate, making execution and customer adoption critical for Weka's differentiation.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No discussion yet for this article
Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime
1 minute a day. The AI essentials.
200+ sources · Email / LINE / Slack