
Frontier AI agents completed the engineering tasks required to conduct AI research but failed to produce publishable findings when tasked with answering open-ended research questions from unpublished academic papers, according to a new evaluation method called shadow evaluations.
The agents lacked judgment about research standards, did not creatively address feedback about flawed designs, and mismanaged resources—burning less than 50% of available compute budgets.
The finding suggests that despite progress on narrow, verifiable tasks, today's agents cannot solve the weeks-long, open-ended research questions that industry leaders claim AI will soon automate.
What happened
Researchers tested two frontier AI agents on unpublished research questions from NeurIPS 2026 papers, giving them six days and $3,000 in API credits each. Despite completing the engineering work without human help, both agents produced papers that original authors rejected unambiguously—scoring 2/6 and 1/6 overall respectively.
Why it matters
Claims that AI will soon automate AI research underpin forecasts of explosive AI progress. This study shows agents can execute research tasks but struggle with judgment, creativity, and resource awareness—the human core of research. The gap between doing experiments and doing publishable research is wider than current benchmarks reveal.
What to watch
The researchers identified five recurring failure modes—poor judgment on research standards, uncreative responses to negative feedback, ineffective backtracking, poor resource awareness, and instruction drift. A robustness check with a second model (GPT-5.6 Sol) and different scaffold reproduced nearly every failure, suggesting the problem is not simply a tool limitation.
Ask the AI about this article →
The research directly challenges a core assumption in forecasts of rapid AI progress: that frontier AI agents can automate AI research. While Anthropic and OpenAI have made public claims about AI systems contributing to research—including OpenAI's statement that GPT-5.6 Sol helped with post-training and saved researchers multiple weeks—the evidence base for agent-driven research capability has been thin. Most prior evaluations focus either on narrow, verifiable tasks with fixed metrics (like optimizing a known benchmark) or on submitting AI-generated papers to peer review, which the authors argue is stochastic and reveals little about how many attempts failed before acceptance.
The shadow evaluation method introduces a third approach: pairing agents with unpublished papers, asking them to answer the same central research question the authors did, and having those authors grade the result. This design avoids contamination (the agent cannot look up published findings) and ensures expert grading on questions they have spent months thinking through. The result is unambiguous: agents can run experiments, write code, and execute a research workflow, but they cannot judge when to stop pursuing a line of inquiry, creatively respond to evidence their approach is failing, or allocate their resources toward the most promising directions. Both agents finished with hours of wall-clock time and less than 50% of their budget remaining, despite papers that did not meet their own success criteria.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
CBTS Technology Solutions LLC launched Forge Agents, a platform that turns a plain-language job description in…
Imec CEO Patrick Vandenameele said at SEMICON Taiwan 2026 that the Belgian research center is broadening its c…

Alphabet's AI Overviews now reach over 2.5 billion monthly users through Google Search, and its ad business ge…

Amazon Web Services (AWS) has integrated its fully managed data warehouse service, Amazon Redshift, with Agent…

Visual Studio Code 1.135 now includes an experimental 'Rubber Duck' feature that lets developers request a sec…

Sonos announced a new app update with generative AI features, a new soundbar called the Beam Ultra, and its se…
