AIToday
Large Language ModelsAI Safety & AlignmentMITテクノロジーレビューPublished: Aug 19, 2026, 10:01 JST3 min read

AI self-improvement hits creativity wall in new research

AI self-improvement hits creativity wall in new research

Key takeaway

  • New research from Princeton and collaborating institutions reveals that while AI agents can perform well-defined engineering tasks in AI research, they lack the judgment and creativity needed to produce original research suitable for publication at top conferences.

  • The study used Claude Opus 4.8 to work on two unpublished NeurIPS 2026 research problems and found that open-ended exploration—deciding which hypotheses to pursue and when to change direction—remains beyond current AI capabilities, suggesting that industry timelines for recursive self-improvement may be overstated.

3 Key Points

  1. What happened

    A research team led by Princeton's Peter Kirilgis and Sayash Kapoor tested whether AI agents can perform open-ended AI research—work without a clear right answer that requires judgment and creativity. They used Claude Opus 4.8 running on OpenClaw to tackle two NeurIPS 2026 research problems over 6 days with $3000 in API credits. The findings: AI agents can handle well-defined engineering tasks, but fall short on the judgment and creativity needed to produce original research publishable at top machine learning conferences.

  2. Why it matters

    The AI industry currently promises that AI will soon improve itself with minimal human oversight through what researchers call recursive self-improvement. But this research suggests that timeline may be overstated. AI agents cannot yet perform the open-ended exploration—choosing which hypotheses to test, deciding what evidence settles a question, knowing when to restart—that is essential to advancing AI research. This gap indicates that some of the timelines claimed for automating AI research may outpace what evidence supports.

  3. What to watch

    The research team introduced a new evaluation method called "shadow evaluation," in which AI agents work on research problems from unpublished high-quality papers, preventing the agents from simply memorizing or finding answers online. This approach may set a new standard for testing whether AI can truly perform open-ended research work.

Ask the AI about this article →

Context & Analysis

The research underscores a critical distinction in AI capability that is often blurred in industry commentary. Existing studies of AI agents automating AI research have largely evaluated performance on bounded tasks—solving engineering problems, post-training small language models, evaluating them on benchmarks—where a correct answer can be verified. These results are real and documented, but they do not address whether AI can perform the exploratory work that actually advances the field. The new evaluation method, shadow evaluation, closes this gap by having AI agents work on unpublished research problems from top-tier conferences, removing the possibility that agents simply memorize or retrieve answers from public sources. The findings align with observations about the difference between tactical execution (writing code, optimizing parameters) and strategic discovery (deciding what problems matter, when evidence is sufficient, when to abandon a direction). This distinction matters precisely because the industry's boldest claims rest on the assumption that AI will soon achieve recursive self-improvement—a process that requires not just executing known tasks but originating new research directions with judgment and taste.

FAQ

What specific AI model did the researchers test?
The research team tested Claude Opus 4.8, made by Anthropic, running on OpenClaw open-source software.
What resources were given to the AI agent?
The agent was given 6 days, $3000 in Anthropic API credits, GPU budget for experiments, a dedicated virtual computer, and access to the open web.
What were the two research problems the AI tackled?
One problem asked whether you can control the "persona" of a large language model by editing its weights (the billions of numbers that encode information learned during training). The other was designing a system to detect when a model's reliability has degraded on predictions based on tabular data.
MITテクノロジーレビューRead Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleIntel, PFN, Toyota explore physical AI and humanoids

The AI news that matters, in one minute each morning.

Sign up free