AIToday
Large Language ModelsAI in HealthcareAI Coding AssistantsTHE DECODERPublished: Aug 29, 2026, 04:00 JST2 min read

Google DeepMind's AI Co-Scientist now runs lab experiments and writes papers

Google DeepMind's AI Co-Scientist now runs lab experiments and writes papers

Key takeaway

  • Google DeepMind's Co-Scientist now plans experiments and runs lab equipment.

  • It validated results in materials science, biology, and computer science.

  • The system cut fabricated key results to 4 percent with reliability modules.

3 Key Points

  1. What happened

    Google DeepMind has expanded its multi-agent system Co-Scientist into a lab-integrated research partner. It now plans experiments, writes code, controls lab equipment, and generates scientific manuscripts, not just hypotheses.

  2. Why it matters

    The system has delivered experimentally validated results across three disciplines—materials science, biology, and computer science. Reliability modules cut fabricated key results down to 4 percent, compared to 90 percent for a comparison system.

  3. What to watch

    In a fully autonomous computer science experiment, the system designed an AI architecture that outperformed six frontier models on health benchmarks after correction, but only showed a significant advantage over the baseline in one of nine categories under human evaluation. The gap between lab assistant and autonomous researcher remains wide.

Ask the AI about this article →

Context & Analysis

The expansion of Co-Scientist from a hypothesis generator to a lab-integrated partner marks a step toward autonomous scientific research. The system's closed-loop workflow—deriving hypotheses, planning experiments, executing code, and generating papers—aims to address the problem of AI fabrication. Prior analyses documented fabrication rates of 80 to 100 percent in existing systems, but Co-Scientist's reliability modules cut key-result fabrication to 4 percent in a double-blind study, though residual issues like selective reporting remain.

The validation across three disciplines shows increasing autonomy, from human-guided synthesis to fully autonomous AI design. However, the computer science experiment revealed a gap between benchmark performance and clinical relevance. Despite outperforming six frontier models on health benchmarks, a human evaluation by three physicians showed only a significant advantage in harm reduction. This disconnect raises questions about what automated benchmarks truly measure.

OpenAI plans to unveil an AI agent this fall for research, highlighting the hype around automated research. Yet the debate persists on whether these systems can discover new knowledge or just make explicit what is in their training data. The researchers acknowledge a long journey ahead before AI can navigate the physical realities of science.

FAQ

How does Co-Scientist reduce fabricated results?
It penalizes fabricated or plagiarized content and uses a verification module that cross-checks every numerical claim against the actual results of executed code. With these modules, fabrication of key results dropped to 4 percent.
What were the outcomes in the three disciplines?
In materials science, Co-Scientist found a safer pathway for a 2D material and synthesized semiconductor films on the first try. In biology, it built an image analysis pipeline that matched unpublished lab results for three out of four shape features. In computer science, it designed an AI architecture that outperformed six frontier models on health benchmarks after correcting for long responses.
Does Co-Scientist require human involvement?
Humans were needed to load samples and precursor materials in materials science, and for refinement after 25 rounds. In the computer science experiment, it ran without any human involvement beyond the initial setup.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Financial data firms race to monetize AI chatbot accessSemafor Tech · 55m ago
  • METR: AI agents colluded, escaped during testSemafor Tech · 55m ago
  • Vibe coding enables self-taught app developmentNikkei AI Stocks · 55m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleEXL acquires AI model developer iMerit