AIToday
Large Language ModelsVentureBeat AIPublished: Sep 2, 2026, 06:00 JST1 min read

Frontier models recall 65% of hidden facts with more thinking

Frontier models recall 65% of hidden facts with more thinking

Key takeaway

  • A new study shows LLMs often know facts but fail to recall them.

  • Frontiers like GPT-5 and Gemini-3 encode 95-98% of tested facts.

  • Thinking longer can recover up to 65% of hidden facts.

3 Key Points

  1. What happened

    A study by Google Research and Technion found that frontier models like GPT-5 and Gemini-3 encode 95-98% of tested facts but fail to surface them during generation. By thinking longer at inference time, these models can recover up to 65% of facts they couldn't directly recall.

  2. Why it matters

    Engineering teams often assume hallucinations mean missing knowledge, leading them to increase model size or expand training data. This research suggests that recall, not encoding, is often the primary bottleneck, so more reliable applications may be built without larger models or external databases.

  3. What to watch

    The study points to inference-time computation as a way to unlock existing knowledge. Future model improvements could focus on better recall strategies rather than just scaling up.

Ask the AI about this article →

Context & Analysis

The study from Google Research and Technion challenges the common assumption that hallucinations stem from missing knowledge. Instead, it finds that models often have the facts encoded parametrically but fail to retrieve them during generation. This suggests that recall, not encoding, is a primary bottleneck for factual accuracy. By demonstrating that inference-time computation—thinking longer—can recover up to 65% of hidden facts, the research implies that engineering teams might improve reliability without necessarily resorting to larger models or external databases. This could shift focus toward optimizing inference processes rather than just scaling up.

FAQ

What is the main finding of the study?
The study found that frontier models like GPT-5 and Gemini-3 encode 95-98% of tested facts but fail to surface them during generation. By thinking longer, they can recover up to 65% of facts they couldn't directly recall.
Why do developers typically think LLMs hallucinate?
Developers typically assume the model lacks the required facts, leading them to increase model size, expand training data, or build complex retrieval architectures.
VentureBeat AIRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Walmart settles opioid claims for $50MTop Companies AI · 3h ago
  • Tim Cook's legacy hinges on Apple's AI betTop Companies AI · 3h ago
  • CrowdStrike Falcon Guardian Targets AI SecurityTop Companies AI · 3h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleOpenAI's Astra reaches 'critical' cyber threshold