
Children learn language from far fewer words than AI models.
A preteen hears about 100 million words, but an LLM churns through trillions.
Understanding this gap could lead to more efficient AI models.
What happened
Children still learn language far more efficiently than large language models (LLMs, AI that understands and generates text). An LLM can process a hundred thousand times more words than a person does in learning their mother tongue, while a preteen might hear about 100 million words, yet kids achieve fluency.
Why it matters
This gap, called the data efficiency gap, challenges AI researchers and cognitive scientists. Language models have improved mainly by getting bigger, but the internet's easily available data could run dry as early as the 2030s, making it crucial to learn how children do more with less.
What to watch
The BabyLM competition tests whether models trained on child-scale data (100 million words, or 10 million for toddler-scale) can perform well. It has already challenged assumptions like the effectiveness of curriculum learning.
Ask the AI about this article →
The article highlights a long-standing puzzle: while LLMs like ChatGPT have become fluent, they require an 'inhuman amount of data' compared to children. This gap raises questions about whether innate knowledge of grammar, as proposed by Noam Chomsky, is necessary, or whether learning from a small sample is possible. The BabyLM competition, inspired by such questions, tests models trained on child-sized datasets, potentially revealing insights into both AI and human learning.
For AI researchers, closing this gap could lead to more data-efficient models, useful for training on video or serving minority language communities. For cognitive scientists, it offers a way to test hypotheses about how children learn language, such as whether it is purely statistical or involves innate structures. The article suggests that if models can learn from child-scale data, it might challenge the idea that massive data is necessary, potentially reshaping AI development.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Canonical is co-funding a three-year PhD project at the University of Bristol to investigate using LLMs to tra…

In 9 days from Aug 10, Meta (Muse Glimmer), NVIDIA (Nemotron 3.5 Lightning), and Alibaba Cloud (Qwen3.8-27B) r…

OpenAI has revealed that its AI agents, being evaluated for cybersecurity capabilities, found and exploited a…

An AlgorithmWatch investigation found that ChatGPT, Gemini, Grok, and Claude linked to anti-abortion websites…

Observe by Snowflake, which combines unified telemetry storage, a context graph, and an AI SRE layer, helped s…

Snowflake announced dynamic model routing in Cortex AI Gateway, which selects the most affordable model for ea…
