
What happened
Apple researchers released DeepAmbigQA, a dataset of 3,600 questions designed to test how well large language models answer complex questions that require both resolving name ambiguity (e.g., multiple films with the same title) and reasoning across multiple steps. Half the questions explicitly test name ambiguity resolution.
Why it matters
Even state-of-the-art models like GPT-5 showed incomplete answers, achieving only 0.13 exact match on ambiguous questions and 0.21 on non-ambiguous questions. The research exposes a gap in how current LLMs handle real-world question answering where they must gather and integrate evidence across large datasets—a critical capability for search-integrated systems.
What to watch
The dataset and automatic generation pipeline (DeepAmbigQAGen) are now available for researchers to benchmark and improve QA systems. The findings underscore the need for more robust models that can reliably produce complete answer sets to complex, ambiguous queries.
Summaries like this, in your inbox every morning.
Open-domain question answering remains a challenging task for large language models, particularly when questions involve ambiguity and require reasoning across multiple pieces of information. The research by Apple scientists addresses a meaningful gap: while existing QA benchmarks test either name ambiguity or multi-hop reasoning in isolation, real-world questions often demand both capabilities simultaneously. By constructing DeepAmbigQA through an automatic pipeline grounded in both text corpora and knowledge graphs, the researchers created a dataset that mirrors the complexity of genuine information-seeking tasks that LLMs with integrated search tools encounter in production.
The performance gap revealed by the experiments is striking and instructive. State-of-the-art models achieving only 0.13 exact match on ambiguous questions versus 0.21 on non-ambiguous ones indicates that name resolution is a significant bottleneck—a finding that has direct implications for any system relying on LLMs to answer factual queries at scale. This distinction matters because incomplete answers in question answering can mislead users or cause them to miss relevant results. The research thus positions itself as a diagnostic tool: rather than proposing a solution, it surfaces the precise nature of the problem, creating a concrete benchmark against which future improvements to QA systems can be measured.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Kalkine Media reports that Home Depot (NYSE:HD) is introducing an AI assistant for its retail operations

NVIDIA announced the NVIDIA Open Agent Safety Platform, an open software platform and reference design to secu…

AMD said Monday it agreed to acquire World Labs, Fei-Fei Li's San Francisco AI lab, for about $8.2 billion in…

NVIDIA announced its NVIDIA Open Agent Safety Platform, combining OpenShell open-source software for secure ag…

Baird said agentic AI will drive Micron to new heights, per a CNBC report

Nvidia unveiled a security platform designed to stop AI agents from going rogue
