
New research from Princeton and collaborating institutions reveals that while AI agents can perform well-defined engineering tasks in AI research, they lack the judgment and creativity needed to produce original research suitable for publication at top conferences.
The study used Claude Opus 4.8 to work on two unpublished NeurIPS 2026 research problems and found that open-ended exploration—deciding which hypotheses to pursue and when to change direction—remains beyond current AI capabilities, suggesting that industry timelines for recursive self-improvement may be overstated.
What happened
A research team led by Princeton's Peter Kirilgis and Sayash Kapoor tested whether AI agents can perform open-ended AI research—work without a clear right answer that requires judgment and creativity. They used Claude Opus 4.8 running on OpenClaw to tackle two NeurIPS 2026 research problems over 6 days with $3000 in API credits. The findings: AI agents can handle well-defined engineering tasks, but fall short on the judgment and creativity needed to produce original research publishable at top machine learning conferences.
Why it matters
The AI industry currently promises that AI will soon improve itself with minimal human oversight through what researchers call recursive self-improvement. But this research suggests that timeline may be overstated. AI agents cannot yet perform the open-ended exploration—choosing which hypotheses to test, deciding what evidence settles a question, knowing when to restart—that is essential to advancing AI research. This gap indicates that some of the timelines claimed for automating AI research may outpace what evidence supports.
What to watch
The research team introduced a new evaluation method called "shadow evaluation," in which AI agents work on research problems from unpublished high-quality papers, preventing the agents from simply memorizing or finding answers online. This approach may set a new standard for testing whether AI can truly perform open-ended research work.
Ask the AI about this article →
The research underscores a critical distinction in AI capability that is often blurred in industry commentary. Existing studies of AI agents automating AI research have largely evaluated performance on bounded tasks—solving engineering problems, post-training small language models, evaluating them on benchmarks—where a correct answer can be verified. These results are real and documented, but they do not address whether AI can perform the exploratory work that actually advances the field. The new evaluation method, shadow evaluation, closes this gap by having AI agents work on unpublished research problems from top-tier conferences, removing the possibility that agents simply memorize or retrieve answers from public sources. The findings align with observations about the difference between tactical execution (writing code, optimizing parameters) and strategic discovery (deciding what problems matter, when evidence is sufficient, when to abandon a direction). This distinction matters precisely because the industry's boldest claims rest on the assumption that AI will soon achieve recursive self-improvement—a process that requires not just executing known tasks but originating new research directions with judgment and taste.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic launched Claude Academy on August 20, a free learning site that explains AI fundamentals and how to…

OpenAI rolled out support on Thursday for controlling Apple's iMessage service via ChatGPT on Mac, enabling th…

As AI technology matures, the bottleneck in the industry is moving beyond semiconductor constraints like GPUs…

OpenAI has launched an Apple Messages plug-in for ChatGPT that lets users connect their Messages inbox to the…

Amazon Bedrock now supports OpenAI GPT-5.6 models (Sol, Terra, and Luna variants) across more than 25 AWS Regi…

Cisco is transforming its digital customer experience (DCX) strategy by embedding AI throughout customer journ…
