
What happened
Lightbits Labs announced general availability of Inferra, a software engine that manages KV cache data across GPU memory and storage to prevent GPU stalls.
Why it matters
Inferra targets neocloud providers and enterprises running large language models with long contexts or many sessions, claiming over 100-fold faster time to first token.
What to watch
The company has production pilots underway and is demonstrating Inferra at the AI Infra Summit in Santa Clara next week.
WHO IT HITSNeocloud providers and enterprises running large language models with long context windows or many simultaneous sessions will be most affected, as Inferra aims to cut time to first token and increase concurrent session density on commodity hardware.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
Lightbits Labs, known for inventing the NVMe over TCP storage protocol, is entering the AI inference market with Inferra, announced in March and now generally available. The engine manages KV cache across GPU high-bandwidth memory, DRAM, and NVMe storage, using predictive prefetching to stage data before the GPU needs it. This disaggregation addresses capacity and retrieval speed limits of GPU memory.
The company claims cache hit rates near 99.9% and positions Inferra as most beneficial for retrieval-augmented generation, AI agents, and heavily shared GPU services. Rasmusson compared the approach to CPU techniques from when processor speeds outpaced memory. Lightbits sees neocloud providers as early adopters due to their utilization and margin pressures.
Inferra's success hinges on whether the claimed benchmark results hold in real-world production. The company has pilots underway and will demonstrate at the AI Infra Summit next week, offering a concrete test of whether its staging predictions can deliver the promised latency improvements for operators.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Michael Burry, famous for The Big Short, is short Nvidia, Palantir and Tesla, and in his Substack newsletter s…

Investors have three creative routes to Anthropic exposure before its expected IPO: buying Alphabet, Amazon, o…

DeepSeek launched V4.1-Flash, a 763B-parameter open-weight model with a causal encoder-decoder architecture

Much of the attention on AI infrastructure buildouts is now tied to sheer compute power, with dominance define…

Barron's reported September 10 that Kepler Computing emerged from stealth with a memory architecture using fer…

Dynatrace acquired Arize AI, adding AI observability, evaluation and agent monitoring to its application obser…