AIToday
Large Language ModelsMIT Technology Review AIPublished: Aug 18, 2026, 22:00 JST3 min read

AI agents fail at creative research, raising doubts on self-improvement timeline

AI agents fail at creative research, raising doubts on self-improvement timeline

Key takeaway

  • A new Princeton-led study found that AI agents cannot yet conduct open-ended research, even though they excel at engineering tasks.

  • Researchers tested Claude Opus 4.8 on unpublished research questions from top-tier machine-learning conference papers, but both resulting papers were rejected for lack of originality and poor judgment.

  • The finding challenges industry hype about rapid recursive self-improvement—where AI systems improve themselves—and suggests that creative, exploratory thinking remains beyond current AI capabilities.

3 Key Points

  1. What happened

    Princeton researchers tested whether AI agents could conduct open-ended research—the kind without clear-cut answers that requires judgment and creativity. They had Claude Opus 4.8 tackle unpublished research questions from NeurIPS 2026 papers over six days with $3,000 in API credits. Both papers were rejected by the original authors; the agents handled engineering tasks well but produced work that lacked novelty and failed to explore ideas properly.

  2. Why it matters

    The AI industry has promised that models will soon improve themselves with minimal human oversight, but this study suggests that milestone is further away than many forecasts claim. AI systems can write code and optimize chips, but they struggle with the open-ended thinking—rethinking approaches, exploring multiple ideas, incorporating feedback—that actual research progress requires. Anthropic cofounder Jack Clark called the lack of creativity a "bearish signal on short recursive self-improvement timelines."

  3. What to watch

    The research team is now testing the same experiment with Mythos, Anthropic's most advanced model launched in April, though it is now available only to approved organizations under Trump administration safety restrictions. The core question remains unresolved: whether recursive self-improvement requires creative breakthroughs or can rely solely on narrower, measurable improvements to speed and benchmarks.

Ask the AI about this article →

Context & Analysis

The promise of recursive self-improvement—where AI systems accelerate their own development—has become the AI industry's boldest near-term milestone. Both OpenAI and Anthropic have publicly charted progress toward this goal. In July, OpenAI highlighted that GPT-5.6 Sol helped post-train a smaller model, saving weeks of work. Anthropic published a blog post in June titled "When AI Builds Itself," signaling confidence in the trajectory. However, the Princeton study exposes a critical gap between what AI agents can do narrowly and what they can do creatively.

The researchers designed "shadow evaluation," a new testing method, to measure open-ended research capabilities—the kind of work that cannot be reduced to yes-or-no answers. Their finding that Claude Opus 4.8 could manage all the engineering legwork (literature review, experiment running, result compilation) but produced papers rejected for lack of originality points to a deeper limitation. Sayash Kapoor, one of the study's leads, attributed this to how AI models are trained: they excel at tasks that can be checked automatically through reinforcement learning, but open-ended work requires training environments that do not yet exist. Anthropic cofounder Jack Clark later echoed this observation, noting the company found similar constraints when attempting to automate AI safety research—a domain also demanding intuitive judgment over rote execution.

The study has limits. It evaluated only two papers, and knowing their graders were human researchers evaluating AI work may have influenced their judgments. The researchers themselves had discretion in study design, introducing potential bias. Yet the findings align with what AI companies report internally and raise a structural question about the path to self-improvement: can AI progress through grinding improvements on narrow, measurable tasks alone, or does recursive self-improvement require the creative leaps—like the invention of transformers—that Kapoor argues have driven the field's biggest breakthroughs? That, Kapoor concludes, is "frankly the trillion-dollar question right now."

FAQ

What exactly did the researchers ask the AI agents to do?
The researchers asked Claude Opus 4.8 to answer research questions from two unpublished papers submitted to NeurIPS 2026—one about controlling a language model's behavior by editing its weights, the other about designing a detector for unreliable spreadsheet-based prediction models. The agents were given six days, $3,000 in Anthropic API credits, a GPU budget, and access to the open web to produce a publishable research paper.
What was the main weakness the agents showed?
The agents struggled with open-ended thinking and creativity. They ran bizarre experiments on tiny synthetic datasets, committed to unpromising approaches too quickly, rejected novel hypotheses based on limited data, and could not fundamentally rethink or restart their methodology. They also could not effectively incorporate feedback or manage resources like tokens, compute time, and paper length.
Why might this matter for claims about AI self-improvement?
The industry has predicted that AI will soon improve itself with minimal human oversight, but this research suggests that capability requires creative, exploratory thinking that current AI systems lack. As Anthropic cofounder Jack Clark noted, the absence of "valuable, intuitive creativity" in today's AI systems is a "bearish signal on short recursive self-improvement timelines."
MIT Technology Review AIRead Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleGoogle's Pet Memory can't tell cats apart after two weeks of testing

The AI news that matters, in one minute each morning.

Sign up free