AIToday
Large Language ModelsAI Business & IndustrySiliconANGLE AIPublished: Sep 10, 2026, 01:00 JST1 min read

Lightbits launches Inferra to boost AI inference performance

Lightbits launches Inferra to boost AI inference performance

3 Key Points

  1. What happened

    Lightbits Labs announced general availability of Inferra, a software engine that manages KV cache data across GPU memory and storage to prevent GPU stalls.

  2. Why it matters

    Inferra targets neocloud providers and enterprises running large language models with long contexts or many sessions, claiming over 100-fold faster time to first token.

  3. What to watch

    The company has production pilots underway and is demonstrating Inferra at the AI Infra Summit in Santa Clara next week.

WHO IT HITSNeocloud providers and enterprises running large language models with long context windows or many simultaneous sessions will be most affected, as Inferra aims to cut time to first token and increase concurrent session density on commodity hardware.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

Lightbits Labs, known for inventing the NVMe over TCP storage protocol, is entering the AI inference market with Inferra, announced in March and now generally available. The engine manages KV cache across GPU high-bandwidth memory, DRAM, and NVMe storage, using predictive prefetching to stage data before the GPU needs it. This disaggregation addresses capacity and retrieval speed limits of GPU memory.

The company claims cache hit rates near 99.9% and positions Inferra as most beneficial for retrieval-augmented generation, AI agents, and heavily shared GPU services. Rasmusson compared the approach to CPU techniques from when processor speeds outpaced memory. Lightbits sees neocloud providers as early adopters due to their utilization and margin pressures.

Inferra's success hinges on whether the claimed benchmark results hold in real-world production. The company has pilots underway and will demonstrate at the AI Infra Summit next week, offering a concrete test of whether its staging predictions can deliver the promised latency improvements for operators.

FAQ
What is KV cache?
KV caches hold intermediate attention data generated as a model processes a prompt. Their size grows with conversation length, consuming GPU memory and causing stalls if unavailable.
Who is the initial target customer for Inferra?
Neocloud providers are the initial target. Lightbits says their need to raise utilization and margins makes them faster adopters than hyperscalers.
What are the claimed performance benefits?
Lightbits claims Inferra cuts time to first token by more than 100-fold in long-context workloads and supports context windows over 10 million tokens on commodity hardware. These are company extrapolations.
SiliconANGLE AIRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • DeepSeek V4.1-Flash: 763B model beats V4 Pro on AA Index 40Latent Space · 1h ago
  • Dynatrace acquires Arize AI as observability shifts to actionSiliconANGLE AI · 7h ago
  • Shared base cuts 100 fine-tunes from 1.5 TB to 19.3 GBDaily Dose of Data Science · 7h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleSystex August revenue up 61.91% on enterprise AI