
Apple researchers introduced DeepAmbigQA, a benchmark of 3,600 questions that test large language models' ability to answer complex queries involving both name ambiguity (like multiple films with the same title) and multi-step reasoning.
Experiments showed that even cutting-edge models like GPT-5 produce incomplete answers, scoring only 0.13 exact match on ambiguous questions and 0.21 on non-ambiguous ones, highlighting a critical weakness in current AI systems designed for open-domain question answering.
What happened
Apple researchers released DeepAmbigQA, a dataset of 3,600 questions designed to test how well large language models answer complex questions that require both resolving name ambiguity (e.g., multiple films with the same title) and reasoning across multiple steps. Half the questions explicitly test name ambiguity resolution.
Why it matters
Even state-of-the-art models like GPT-5 showed incomplete answers, achieving only 0.13 exact match on ambiguous questions and 0.21 on non-ambiguous questions. The research exposes a gap in how current LLMs handle real-world question answering where they must gather and integrate evidence across large datasets—a critical capability for search-integrated systems.
What to watch
The dataset and automatic generation pipeline (DeepAmbigQAGen) are now available for researchers to benchmark and improve QA systems. The findings underscore the need for more robust models that can reliably produce complete answer sets to complex, ambiguous queries.
Apple researchers have developed DeepAmbigQA, a new benchmark that exposes a critical limitation in how large language models handle complex question answering. The work, authored by Jiabao Ji, Min Li, Priyanshu Kumar, Shiyu Chang, and Saloni Potdar, focuses on questions that require both resolving name ambiguity and reasoning across multiple pieces of information—two challenges that rarely appear together in existing benchmarks despite being common in real-world use.
The core motivation stems from a practical observation: LLMs equipped with search tools often fail on questions like "Which actor from the film Heat won at least one Academy Award?" This question demands two distinct cognitive tasks. First, the system must recognize that "Heat" refers to a specific film among possibly several with the same title. Second, it must reason across the cast of that film and cross-reference Academy Award winners to compile a complete answer. The researchers built DeepAmbigQA using an automatic data generation pipeline called DeepAmbigQAGen, which grounds QA tasks in both text corpora and linked knowledge graphs to ensure questions are natural, verifiable, and systematically embed both name ambiguity and multi-step reasoning requirements.
The resulting dataset contains 3,600 questions, with exactly half explicitly requiring name ambiguity resolution. When tested against state-of-the-art models, the results reveal a substantial gap: GPT-5 achieved only 0.13 exact match on ambiguous questions and 0.21 on non-ambiguous questions. This performance differential underscores that name disambiguation is a harder problem than the non-ambiguous variants, and that even the most advanced models struggle to produce complete answer sets. The findings highlight that more robust QA systems are needed—systems specifically engineered to handle information gathering and answer completeness in the presence of ambiguity.
Open-domain question answering remains a challenging task for large language models, particularly when questions involve ambiguity and require reasoning across multiple pieces of information. The research by Apple scientists addresses a meaningful gap: while existing QA benchmarks test either name ambiguity or multi-hop reasoning in isolation, real-world questions often demand both capabilities simultaneously. By constructing DeepAmbigQA through an automatic pipeline grounded in both text corpora and knowledge graphs, the researchers created a dataset that mirrors the complexity of genuine information-seeking tasks that LLMs with integrated search tools encounter in production.
The performance gap revealed by the experiments is striking and instructive. State-of-the-art models achieving only 0.13 exact match on ambiguous questions versus 0.21 on non-ambiguous ones indicates that name resolution is a significant bottleneck—a finding that has direct implications for any system relying on LLMs to answer factual queries at scale. This distinction matters because incomplete answers in question answering can mislead users or cause them to miss relevant results. The research thus positions itself as a diagnostic tool: rather than proposing a solution, it surfaces the precise nature of the problem, creating a concrete benchmark against which future improvements to QA systems can be measured.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Amazon and Google are intensifying competitive efforts against The Trade Desk (TTD), a major digital advertisi…

OpenAI introduced Premium Seats for ChatGPT Business, priced at $125 per user per month ($100 with annual bill…

Computer scientists at University of Tübingen, Max Planck Institute, MATS Research, and Snyk discovered a meth…

Anthropic pledged to embed machine-readable watermarks in Claude-generated text and digitally signed provenanc…

Anthropic has signed the EU AI Act Code of Practice and will embed invisible watermarks in Claude-generated te…

Anthropic has agreed to pay $9.1 billion over 20 years to Riot Platforms Inc., a Bitcoin miner turned data cen…

The AI news that matters, in one minute each morning.
Sign up free