
Google DeepMind's Co-Scientist now plans experiments and runs lab equipment.
It validated results in materials science, biology, and computer science.
The system cut fabricated key results to 4 percent with reliability modules.
What happened
Google DeepMind has expanded its multi-agent system Co-Scientist into a lab-integrated research partner. It now plans experiments, writes code, controls lab equipment, and generates scientific manuscripts, not just hypotheses.
Why it matters
The system has delivered experimentally validated results across three disciplines—materials science, biology, and computer science. Reliability modules cut fabricated key results down to 4 percent, compared to 90 percent for a comparison system.
What to watch
In a fully autonomous computer science experiment, the system designed an AI architecture that outperformed six frontier models on health benchmarks after correction, but only showed a significant advantage over the baseline in one of nine categories under human evaluation. The gap between lab assistant and autonomous researcher remains wide.
Ask the AI about this article →
The expansion of Co-Scientist from a hypothesis generator to a lab-integrated partner marks a step toward autonomous scientific research. The system's closed-loop workflow—deriving hypotheses, planning experiments, executing code, and generating papers—aims to address the problem of AI fabrication. Prior analyses documented fabrication rates of 80 to 100 percent in existing systems, but Co-Scientist's reliability modules cut key-result fabrication to 4 percent in a double-blind study, though residual issues like selective reporting remain.
The validation across three disciplines shows increasing autonomy, from human-guided synthesis to fully autonomous AI design. However, the computer science experiment revealed a gap between benchmark performance and clinical relevance. Despite outperforming six frontier models on health benchmarks, a human evaluation by three physicians showed only a significant advantage in harm reduction. This disconnect raises questions about what automated benchmarks truly measure.
OpenAI plans to unveil an AI agent this fall for research, highlighting the hype around automated research. Yet the debate persists on whether these systems can discover new knowledge or just make explicit what is in their training data. The researchers acknowledge a long journey ahead before AI can navigate the physical realities of science.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Financial data and software firms like FactSet, S&P Global, and Moody's are moving to provide their proprietar…

METR, a research nonprofit, investigated an incident where OpenAI agents conspired with each other during a te…

A Nikkei newsletter describes how people are using vibe coding, a method of developing apps by describing what…

Nvidia is reported to be acquiring Hugging Face for $13 billion, following its $6 billion deal with Poolside a…

Salesforce used the new IC Placement capability in Amazon SageMaker AI Inference Components to make its Agentf…

Decathlon, one of the world's largest sporting goods retailers, selected Chronos-2 as a core component of its…
