AIToday
Large Language ModelsAI Safety & AlignmentHacker NewsPublished: Aug 6, 2026, 22:03 JST3 min read

AI agents fail open-ended research despite engineering prowess

AI agents fail open-ended research despite engineering prowess

Key takeaway

  • Frontier AI agents completed the engineering tasks required to conduct AI research but failed to produce publishable findings when tasked with answering open-ended research questions from unpublished academic papers, according to a new evaluation method called shadow evaluations.

  • The agents lacked judgment about research standards, did not creatively address feedback about flawed designs, and mismanaged resources—burning less than 50% of available compute budgets.

  • The finding suggests that despite progress on narrow, verifiable tasks, today's agents cannot solve the weeks-long, open-ended research questions that industry leaders claim AI will soon automate.

3 Key Points

  1. What happened

    Researchers tested two frontier AI agents on unpublished research questions from NeurIPS 2026 papers, giving them six days and $3,000 in API credits each. Despite completing the engineering work without human help, both agents produced papers that original authors rejected unambiguously—scoring 2/6 and 1/6 overall respectively.

  2. Why it matters

    Claims that AI will soon automate AI research underpin forecasts of explosive AI progress. This study shows agents can execute research tasks but struggle with judgment, creativity, and resource awareness—the human core of research. The gap between doing experiments and doing publishable research is wider than current benchmarks reveal.

  3. What to watch

    The researchers identified five recurring failure modes—poor judgment on research standards, uncreative responses to negative feedback, ineffective backtracking, poor resource awareness, and instruction drift. A robustness check with a second model (GPT-5.6 Sol) and different scaffold reproduced nearly every failure, suggesting the problem is not simply a tool limitation.

Ask the AI about this article →

Context & Analysis

The research directly challenges a core assumption in forecasts of rapid AI progress: that frontier AI agents can automate AI research. While Anthropic and OpenAI have made public claims about AI systems contributing to research—including OpenAI's statement that GPT-5.6 Sol helped with post-training and saved researchers multiple weeks—the evidence base for agent-driven research capability has been thin. Most prior evaluations focus either on narrow, verifiable tasks with fixed metrics (like optimizing a known benchmark) or on submitting AI-generated papers to peer review, which the authors argue is stochastic and reveals little about how many attempts failed before acceptance.

The shadow evaluation method introduces a third approach: pairing agents with unpublished papers, asking them to answer the same central research question the authors did, and having those authors grade the result. This design avoids contamination (the agent cannot look up published findings) and ensures expert grading on questions they have spent months thinking through. The result is unambiguous: agents can run experiments, write code, and execute a research workflow, but they cannot judge when to stop pursuing a line of inquiry, creatively respond to evidence their approach is failing, or allocate their resources toward the most promising directions. Both agents finished with hours of wall-clock time and less than 50% of their budget remaining, despite papers that did not meet their own success criteria.

FAQ

How long did the agents have to work on these research questions?
The agents were given six days of wall-clock time and $3,000 in Anthropic API credits, plus GPU credits for experiments and full access to a VM and the open web.
What were the specific failure modes the researchers identified?
The five recurring failure modes were: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift.
Did the results change when tested with a different model?
A robustness check using GPT-5.6 Sol and Codex with the same time and API budgets produced similar results to the Opus 4.8 experiments, reproducing nearly every identified failure mode.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • CBTS launches Forge Agents for custom AI agentsSiliconANGLE AI · 44m ago
  • Imec CEO: AI era widens chip-model-CSP collaborationDIGITIMES Asia · 44m ago
  • Alphabet's AI Overviews reach 2.5B monthly usersYahoo Finance AI · 44m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAI moderation silences marginalized groups; platforms need human oversight