AIToday
Large Language ModelsApple Machine LearningPublished: Aug 7, 2026, 10:01 JST

Apple researchers reveal LLMs struggle with ambiguous multi-step questions

Apple researchers reveal LLMs struggle with ambiguous multi-step questions

3 Key Points

  1. What happened

    Apple researchers released DeepAmbigQA, a dataset of 3,600 questions designed to test how well large language models answer complex questions that require both resolving name ambiguity (e.g., multiple films with the same title) and reasoning across multiple steps. Half the questions explicitly test name ambiguity resolution.

  2. Why it matters

    Even state-of-the-art models like GPT-5 showed incomplete answers, achieving only 0.13 exact match on ambiguous questions and 0.21 on non-ambiguous questions. The research exposes a gap in how current LLMs handle real-world question answering where they must gather and integrate evidence across large datasets—a critical capability for search-integrated systems.

  3. What to watch

    The dataset and automatic generation pipeline (DeepAmbigQAGen) are now available for researchers to benchmark and improve QA systems. The findings underscore the need for more robust models that can reliably produce complete answer sets to complex, ambiguous queries.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Open-domain question answering remains a challenging task for large language models, particularly when questions involve ambiguity and require reasoning across multiple pieces of information. The research by Apple scientists addresses a meaningful gap: while existing QA benchmarks test either name ambiguity or multi-hop reasoning in isolation, real-world questions often demand both capabilities simultaneously. By constructing DeepAmbigQA through an automatic pipeline grounded in both text corpora and knowledge graphs, the researchers created a dataset that mirrors the complexity of genuine information-seeking tasks that LLMs with integrated search tools encounter in production.

The performance gap revealed by the experiments is striking and instructive. State-of-the-art models achieving only 0.13 exact match on ambiguous questions versus 0.21 on non-ambiguous ones indicates that name resolution is a significant bottleneck—a finding that has direct implications for any system relying on LLMs to answer factual queries at scale. This distinction matters because incomplete answers in question answering can mislead users or cause them to miss relevant results. The research thus positions itself as a diagnostic tool: rather than proposing a solution, it surfaces the precise nature of the problem, creating a concrete benchmark against which future improvements to QA systems can be measured.

FAQ
What is DeepAmbigQA and how large is the dataset?
DeepAmbigQA is a dataset of 3,600 questions requiring multi-hop reasoning, with half of them explicitly testing name ambiguity resolution. It was built using DeepAmbigQAGen, an automatic data generation pipeline that constructs QA tasks grounded in text corpora and linked knowledge graphs.
How well do current state-of-the-art LLMs perform on this benchmark?
Even GPT-5 achieved only 0.13 exact match on ambiguous questions and 0.21 on non-ambiguous questions, demonstrating that current models struggle to produce complete answer sets to complex questions requiring both name disambiguation and multi-step reasoning.
What type of questions does the benchmark test?
The benchmark tests questions that require distinguishing between multiple entities sharing the same name (such as films with identical titles) and reasoning across large sets of entities to gather and integrate evidence, such as 'Which actor from the film Heat won at least one Academy Award?'
Apple Machine LearningRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Nvidia unveils security platform to stop rogue AI agentsTop Companies AI · 43m ago
  • AMD to buy Fei-Fei Li's World Labs for $8.2 billionTop Companies AI · 43m ago
  • Home Depot (NYSE:HD) rolls out AI assistant for shoppersTop Companies AI · 43m ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleAmphenol's AI Infrastructure Push Valued at 17.9x Forward EV/EBITDA