AIToday
Large Language ModelsVentureBeat AIPublished: Mar 28, 2026, 04:00 JST1 min read

IndexCache technique from Tsinghua University and Z.ai accelerates long-context LLM inference by up to 1.82x by eliminating redundant sparse attention computations.

3 Key Points

  1. IndexCache reduces redundant computation in sparse attention models by up to 75%, significantly speeding up processing of 200,000-token contexts

  2. Delivers 1.82x faster time-to-first-token and 1.48x faster generation throughput for long-context inference tasks

  3. Compatible with DeepSeek Sparse Attention architecture, including DeepSeek and GLM model families

  4. Successfully tested on GLM-5, a 744-billion-parameter model, proving effectiveness at production scale

  5. Addresses the bottleneck of self-attention mechanisms in LLMs, enabling faster enterprise-grade long-context experiences

Ask the AI about this article →

VentureBeat AIRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • DataAgent launches with $10M to auto-fix Kubernetes faultsSiliconANGLE AI · 1h ago
  • SK Hynix custom HBM boosts inference up to 5.15xDIGITIMES Asia · 1h ago
  • Nvidia Earnings: Boring by Design, Avoiding a Consolidated WorldStratechery (Ben Thompson) · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleTech companies face mounting pressure over data center energy demands as senators investigate power grid impacts and environmental concerns