AIToday
Large Language ModelsApple Machine LearningPublished: Aug 7, 2026, 10:01 JST4 min read

Apple researchers reveal LLMs struggle with ambiguous multi-step questions

Apple researchers reveal LLMs struggle with ambiguous multi-step questions

Key takeaway

  • Apple researchers introduced DeepAmbigQA, a benchmark of 3,600 questions that test large language models' ability to answer complex queries involving both name ambiguity (like multiple films with the same title) and multi-step reasoning.

  • Experiments showed that even cutting-edge models like GPT-5 produce incomplete answers, scoring only 0.13 exact match on ambiguous questions and 0.21 on non-ambiguous ones, highlighting a critical weakness in current AI systems designed for open-domain question answering.

3 Key Points

  1. What happened

    Apple researchers released DeepAmbigQA, a dataset of 3,600 questions designed to test how well large language models answer complex questions that require both resolving name ambiguity (e.g., multiple films with the same title) and reasoning across multiple steps. Half the questions explicitly test name ambiguity resolution.

  2. Why it matters

    Even state-of-the-art models like GPT-5 showed incomplete answers, achieving only 0.13 exact match on ambiguous questions and 0.21 on non-ambiguous questions. The research exposes a gap in how current LLMs handle real-world question answering where they must gather and integrate evidence across large datasets—a critical capability for search-integrated systems.

  3. What to watch

    The dataset and automatic generation pipeline (DeepAmbigQAGen) are now available for researchers to benchmark and improve QA systems. The findings underscore the need for more robust models that can reliably produce complete answer sets to complex, ambiguous queries.

In Depth

Read the full story

Apple researchers have developed DeepAmbigQA, a new benchmark that exposes a critical limitation in how large language models handle complex question answering. The work, authored by Jiabao Ji, Min Li, Priyanshu Kumar, Shiyu Chang, and Saloni Potdar, focuses on questions that require both resolving name ambiguity and reasoning across multiple pieces of information—two challenges that rarely appear together in existing benchmarks despite being common in real-world use.

The core motivation stems from a practical observation: LLMs equipped with search tools often fail on questions like "Which actor from the film Heat won at least one Academy Award?" This question demands two distinct cognitive tasks. First, the system must recognize that "Heat" refers to a specific film among possibly several with the same title. Second, it must reason across the cast of that film and cross-reference Academy Award winners to compile a complete answer. The researchers built DeepAmbigQA using an automatic data generation pipeline called DeepAmbigQAGen, which grounds QA tasks in both text corpora and linked knowledge graphs to ensure questions are natural, verifiable, and systematically embed both name ambiguity and multi-step reasoning requirements.

The resulting dataset contains 3,600 questions, with exactly half explicitly requiring name ambiguity resolution. When tested against state-of-the-art models, the results reveal a substantial gap: GPT-5 achieved only 0.13 exact match on ambiguous questions and 0.21 on non-ambiguous questions. This performance differential underscores that name disambiguation is a harder problem than the non-ambiguous variants, and that even the most advanced models struggle to produce complete answer sets. The findings highlight that more robust QA systems are needed—systems specifically engineered to handle information gathering and answer completeness in the presence of ambiguity.

Context & Analysis

Open-domain question answering remains a challenging task for large language models, particularly when questions involve ambiguity and require reasoning across multiple pieces of information. The research by Apple scientists addresses a meaningful gap: while existing QA benchmarks test either name ambiguity or multi-hop reasoning in isolation, real-world questions often demand both capabilities simultaneously. By constructing DeepAmbigQA through an automatic pipeline grounded in both text corpora and knowledge graphs, the researchers created a dataset that mirrors the complexity of genuine information-seeking tasks that LLMs with integrated search tools encounter in production.

The performance gap revealed by the experiments is striking and instructive. State-of-the-art models achieving only 0.13 exact match on ambiguous questions versus 0.21 on non-ambiguous ones indicates that name resolution is a significant bottleneck—a finding that has direct implications for any system relying on LLMs to answer factual queries at scale. This distinction matters because incomplete answers in question answering can mislead users or cause them to miss relevant results. The research thus positions itself as a diagnostic tool: rather than proposing a solution, it surfaces the precise nature of the problem, creating a concrete benchmark against which future improvements to QA systems can be measured.

FAQ

What is DeepAmbigQA and how large is the dataset?
DeepAmbigQA is a dataset of 3,600 questions requiring multi-hop reasoning, with half of them explicitly testing name ambiguity resolution. It was built using DeepAmbigQAGen, an automatic data generation pipeline that constructs QA tasks grounded in text corpora and linked knowledge graphs.
How well do current state-of-the-art LLMs perform on this benchmark?
Even GPT-5 achieved only 0.13 exact match on ambiguous questions and 0.21 on non-ambiguous questions, demonstrating that current models struggle to produce complete answer sets to complex questions requiring both name disambiguation and multi-step reasoning.
What type of questions does the benchmark test?
The benchmark tests questions that require distinguishing between multiple entities sharing the same name (such as films with identical titles) and reasoning across large sets of entities to gather and integrate evidence, such as 'Which actor from the film Heat won at least one Academy Award?'
Apple Machine LearningRead Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAstera Labs, Amphenol emerge as AI infrastructure profit play

The AI news that matters, in one minute each morning.

Sign up free