AIToday
Large Language ModelsMITテクノロジーレビューPublished: Aug 28, 2026, 10:01 JST1 min read

Why Children Learn Language Faster Than AI

Why Children Learn Language Faster Than AI

Key takeaway

  • Children learn language with far less data than AI models.

  • This gap matters because internet data may run out.

  • Reverse-engineering child learning could make AI more efficient.

3 Key Points

  1. What happened

    A new analysis highlights the 'data efficiency gap'—children learn language with far less data than AI models. For example, an average child hears about 100 million words by age 10, while LLMs like Llama 3.1 process 15 trillion tokens.

  2. Why it matters

    This gap poses a challenge for AI designers as readily available internet data may run out by the early 2030s. Understanding how children learn could lead to more data-efficient AI models, useful for video-based learning or chatbots for minority language communities.

  3. What to watch

    Researchers are reverse-engineering children's learning to improve AI. This could also settle long-standing questions about whether language is innate or learned purely from experience.

Ask the AI about this article →

Context & Analysis

The article traces decades of AI language research, from Chomsky's theory of innate grammar to the rise of statistical models like BERT and GPT-2. The shift to data-hungry neural networks solved fluency but at the cost of efficiency. This efficiency gap is not just academic; it has practical implications as data supply dwindles.

Researchers hope that studying child development can inspire new AI architectures that learn from less data, potentially enabling applications like video-based learning or supporting minority languages. The stakes include settling debates at the intersection of linguistics and cognitive science—questions about whether language is innate or learned, and how biology shapes our language abilities.

The commentary notes that while LLMs are powerful statistical learners, they are not brains. The challenge is to close the efficiency gap without sacrificing capability, a goal that remains unrealized but promising.

FAQ

How much data does a child need to learn language?
A child typically starts speaking grammatically after hearing 10 to 30 million words, and by age 10 may have heard around 100 million words.
Why is this an issue for AI?
LLMs require vast amounts of data—like Llama 3.1's 15 trillion tokens—and available internet data may be exhausted by the early 2030s.
MITテクノロジーレビューRead Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleHonor Robot Phone enters mass production