AIToday
Large Language ModelsTHE DECODERPublished: Aug 20, 2026, 22:00 JST3 min read

Sutton: synthetic data a 'big mistake' for scaling AI

Sutton: synthetic data a 'big mistake' for scaling AI

Key takeaway

  • Richard Sutton, a pioneering AI researcher and Turing Award winner, has publicly rejected synthetic data as a path forward for scaling large language models, calling it 'a big mistake.' He argues the world is infinitely complex—far larger than any simulation humans can build—and that synthetic data strategies create an unavoidable human bottleneck.

  • Instead, Sutton advocates for agents that learn continuously from their own experience and adjust their world models in real time, a direction he says addresses a fundamental limitation of today's language models, which he views as only 'like 20% or a quarter of intelligence.'

3 Key Points

  1. What happened

    Turing Award winner Richard Sutton, a founder of reinforcement learning, argued that synthetic data cannot solve the scaling bottleneck for large language models. Sutton, who recently founded Oak Lab with former student Khurram Javeed, says the world is infinitely complex and any simulation of it is 'microscopic' compared to reality.

  2. Why it matters

    Current LLMs have exhausted the finite internet as a training source, leaving AI labs searching for new data. Sutton contends that synthetic data strategies create a human bottleneck—domain experts must judge which data is good or bad—and cannot scale without human oversight. His critique challenges a core strategy at leading AI labs.

  3. What to watch

    Sutton proposes an alternative: agents should learn continuously from their own experience and build their own world models, rather than rely on frozen human-built simulations. He points to 'Continual Backprop,' a method his team published in Nature, as a step toward real continual learning; he views current language models as only 'like 20% or a quarter of intelligence.'

Ask the AI about this article →

Context & Analysis

Richard Sutton's critique of synthetic data strategies strikes at a central pivot point in AI labs' efforts to extend training beyond internet-scale data. With large language models now constrained by the finiteness of publicly available text, companies including Anthropic, OpenAI, and others have increasingly turned to synthetic data—model-generated examples designed to augment training sets. Sutton's objection rests on a foundational claim: the "Big World Hypothesis," formulated by his cofounder Khurram Javeed and developed over years at the University of Alberta, posits that any deterministic simulation of the world will be orders of magnitude simpler than reality itself and thus fundamentally mismatched to it.

The practical implication Sutton draws is that domain expertise becomes a hard constraint. Whether training a drone to move like a bat using echolocation or refining self-driving car behavior, humans must bridge the gap between simulation and lived reality—and that human involvement cannot be scaled away. This argument directly challenges the efficiency narrative behind synthetic data: it does not eliminate the need for human curation; it merely relocates it. Sutton's alternative—continuous learning from an agent's own experience, adjusting a self-built world model—sidesteps the human bottleneck by removing the need for pre-built simulations altogether. His framing of current language models as only "like 20% or a quarter of intelligence" suggests he sees this limitation as fundamental to the training paradigm itself, not a temporary obstacle to be solved with more data.

FAQ

Why does Sutton think synthetic data won't work for training AI?
Sutton argues the world is infinitely complex and 'massively more complex than your mind.' Any simulation is 'microscopic' by comparison, with wrong friction values or inaccurate models. Additionally, the world contains other agents whose inner workings cannot be generated as synthetic data—'there's no way we can have synthetic data for other people's minds.' A second problem is the human bottleneck: domain experts must decide which synthetic data is good or bad, and this requirement for human expertise prevents scaling.
What is Sutton's alternative to synthetic data?
Sutton proposes letting agents learn from their own experience instead. An agent should learn its own world model and continuously correct it, rather than rely on a frozen simulation model built by humans. He emphasizes real continual learning—learning without forgetting old knowledge—and points to 'Continual Backprop,' a method his team published in Nature, as a step toward this approach.
Who is Richard Sutton and what is Oak Lab?
Sutton is a Turing Award winner and one of the founders of reinforcement learning; he wrote the field's standard textbook and authored the influential 2019 essay 'The Bitter Lesson.' He recently founded Oak Lab with his former student Khurram Javeed. The company focuses on the 'Big World Hypothesis'—the assumption that the world is infinitely complex and cannot be adequately simulated.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleBond yields hit pre-2008 crisis levels; Treasury doubles buyback plan

The AI news that matters, in one minute each morning.

Sign up free