AIToday
Large Language ModelsMIT Technology Review AIPublished: Aug 24, 2026, 19:00 JST2 min read

Children vs. AI: the data efficiency gap

Children vs. AI: the data efficiency gap

Key takeaway

  • Children learn language from far fewer words than AI models.

  • A preteen hears about 100 million words, but an LLM churns through trillions.

  • Understanding this gap could lead to more efficient AI models.

3 Key Points

  1. What happened

    Children still learn language far more efficiently than large language models (LLMs, AI that understands and generates text). An LLM can process a hundred thousand times more words than a person does in learning their mother tongue, while a preteen might hear about 100 million words, yet kids achieve fluency.

  2. Why it matters

    This gap, called the data efficiency gap, challenges AI researchers and cognitive scientists. Language models have improved mainly by getting bigger, but the internet's easily available data could run dry as early as the 2030s, making it crucial to learn how children do more with less.

  3. What to watch

    The BabyLM competition tests whether models trained on child-scale data (100 million words, or 10 million for toddler-scale) can perform well. It has already challenged assumptions like the effectiveness of curriculum learning.

Ask the AI about this article →

Context & Analysis

The article highlights a long-standing puzzle: while LLMs like ChatGPT have become fluent, they require an 'inhuman amount of data' compared to children. This gap raises questions about whether innate knowledge of grammar, as proposed by Noam Chomsky, is necessary, or whether learning from a small sample is possible. The BabyLM competition, inspired by such questions, tests models trained on child-sized datasets, potentially revealing insights into both AI and human learning.

For AI researchers, closing this gap could lead to more data-efficient models, useful for training on video or serving minority language communities. For cognitive scientists, it offers a way to test hypotheses about how children learn language, such as whether it is purely statistical or involves innate structures. The article suggests that if models can learn from child-scale data, it might challenge the idea that massive data is necessary, potentially reshaping AI development.

FAQ

What is the data efficiency gap?
It is the divide between the huge amounts of data LLMs need (a hundred thousand times more words than a person) and the relatively small amount children need to learn a language fluently.
How does BabyLM address this gap?
BabyLM is an annual competition where researchers train language models on a 'developmentally plausible' corpus of just 100 million words (or 10 million for the toddler-scale track). The models are evaluated on grammar benchmarks similar to those used with humans, and it has already challenged some assumptions, like the effectiveness of curriculum learning.
MIT Technology Review AIRead Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleNvidia reportedly in talks to invest in Perplexity