
What happened
Oriol Vinyals, until recently VP of Research at Google DeepMind, told the Agentic AI Summit 2026 that recursive self-improvement is coming but slow, with no intelligence explosion in sight.
Why it matters
AI can already speed up implementation and experimentation, but Vinyals says idea generation (research taste) and reliable evaluation still fall short — he expects it to take more time.
What to watch
Vinyals is launching Discovery Loop, co-founded with Jeff Dean as CEO, Sanjay Ghemawat, and Quoc Le, and plans to automate AI research first with the startup as its own first customer. Whether it clears those two bottlenecks is the test.
WHO IT HITSAI lab leaders and research teams planning to lean on automated research pipelines may need to temper expectations; Vinyals's argument suggests the returns hinge on progress in idea generation and evaluation, not raw compute. Founders and investors backing research-automation startups could treat Discovery Loop as a bellwether.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
Vinyals's talk came days after he left Google DeepMind, where he served as VP of Research and worked on AlphaStar, AlphaCode, and Gemini. His argument is less a prediction than a diagnosis of where AI self-improvement stalls. Labs today mostly infer self-improvement indirectly from capability benchmarks like SWE-Bench Pro or ML-Bench, but those tests mainly cover implementation and experimentation — the steps Vinyals says already work. More direct benchmarks are starting to appear, yet they are expensive: each evaluation requires an agent to work for hours on tasks far removed from the real goal, the way optimizing Tetris is not the same as automating a research lab.
Idea generation is equally underdeveloped. Vinyals calls the instinct for which ideas are worth pursuing "research taste," and notes nobody has really studied how to teach it in LLM training. He expects future evaluations to reward not just how much a system improves but how it gets there, using the criteria conference reviewers apply: originality, elegance, efficiency, and durability. Human review is expensive and not especially good at spotting strong ideas either.
Discovery Loop, his new startup, aims to automate the full scientific loop — hypotheses, experiments, evaluation — starting with AI research itself, with the company as its own first customer. Vinyals concedes idea generation remains the hardest part, so early on humans and machines will develop hypotheses together. Whether that hybrid approach actually moves the needle on the two bottlenecks is the open question, and the founders' stated ambition — that a handful of people could outproduce massive teams — rests on it.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
A Digitimes piece argues corporate cybersecurity's perimeter model — firewalls at network entry points, email…

Much of the attention on AI infrastructure buildouts is now tied to sheer compute power, with dominance define…

Barron's reported September 10 that Kepler Computing emerged from stealth with a memory architecture using fer…

Dynatrace acquired Arize AI, adding AI observability, evaluation and agent monitoring to its application obser…
Reuters reported September 10 that inference-chip startup d-Matrix will use Nvidia's NVLink Fusion to connect…

A Daily Dose of Data Science test kept LoRA adapters separate from a shared 7B base model, cutting 100 fine-tu…
