AIToday
Large Language ModelsQiita 機械学習Published: Oct 6, 2026, 16:00 JST

DeepMind: Chinchilla at 70 billion parameters beats Gopher at 280 billion

DeepMind: Chinchilla at 70 billion parameters beats Gopher at 280 billion

3 Key Points

  1. What happened

    DeepMind's 2022 study found that at equal compute, balancing model size and data beats pure bigness: a 70 billion parameter Chinchilla trained on 1.3 trillion tokens outperformed the 280 billion parameter Gopher, which read about 300 billion tokens.

  2. Why it matters

    The result challenged the long-held scaling rule that predicted gains simply from making models larger, suggesting the bigger model is not automatically the better one when compute is held equal.

  3. What to watch

    The body notes the roughly 20-times-tokens-to-parameters guideline is only a cost-minimizing benchmark, not a formula for the strongest model, and Meta reported Llama 3 kept improving past it, so what real builders do may diverge from it.

WHO IT HITSAI research teams planning training runs — deciding how to split a compute budget between model size and data — now have evidence that the data side can matter more than raw parameter count, which may shift how they scope large training projects.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The article frames this result as a correction to a long-standing assumption. For years, the working rule known as the scaling law held that adding model size, data, and compute reliably improved performance. The 2022 DeepMind study tested that assumption directly by holding compute constant and varying how it was spent.

The outcome — a 70 billion parameter model beating a 280 billion parameter one — suggested the field had been over-investing in size relative to data. The article places this alongside other shifts in how modern AI systems are built, from human feedback and automated grading to efficiency techniques and system design, as part of a broader move away from a single scaling lever.

What the result ultimately means for builders hinges on how far the balance principle generalizes. The article notes the widely cited 20-times-tokens-to-parameters ratio is a cost-minimizing guideline rather than a recipe for the strongest model, and that Meta reported continued gains with far more data. It is likely the field will keep testing where the balance point sits.

FAQ
What did DeepMind's 2022 study actually find?
At equal compute, a 70 billion parameter Chinchilla trained on 1.3 trillion tokens outperformed the 280 billion parameter Gopher, which read about 300 billion tokens, suggesting balance between size and data matters more than sheer scale.
Did the scaling rule change after this study?
The article says the study revised the long-held scaling rule: rather than simply making models bigger, the result favors growing size and training data in balance for the same compute.
Qiita 機械学習Read Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleFalcon-Emirati-7B: 84.83% Alyah score tops Arabic AI rivals