AIToday
Large Language ModelsQiita 機械学習Published: Oct 4, 2026, 22:00 JST

Ouro to 7.7T tokens: Looped Transformers, explained

Ouro to 7.7T tokens: Looped Transformers, explained

3 Key Points

  1. What happened

    A Qiita review traces Looped Transformers from Universal Transformers in 2018 through Giannou et al.'s 2023 programmable-computer paper to ByteDance Seed's Ouro, which trained 1.4B and 2.6B parameter models on up to 7.7T tokens.

  2. Why it matters

    The review frames the same-block repetition not as mere weight sharing but as a way to treat depth as repeated algorithm steps and to separate parameter count from effective depth.

  3. What to watch

    Whether models can learn how many loops a given problem needs; the review notes that a model trained at 8 loops may not improve at 100 loops, and that simple halting gates are not always optimal.

WHO IT HITSThis summary lands on engineers and technical decision-makers tracking efficient LLM architectures, and on researchers evaluating whether parameter savings actually translate into latency or reasoning gains.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The review comes from an author who admits he had filed Looped Transformers under 'just passing a Transformer through the same layers a few times' and stopped there. Following the papers, he argues, does not end at that description. Universal Transformers in 2018 already applied self-attention recursively along depth and added Adaptive Computation Time to stop computation per token — an early version of the current idea. What changed is the framing: Giannou et al. treated the loop not as a way to shrink a model but as a device for executing an algorithm, and Yang et al. tested whether ordinary training would recover repeated-algorithm behavior. The review also notes the input-injection trick added in that later work, where the original input is re-added at each loop so early information is not buried in the hidden state.

The later papers push scale and adaptivity. A 3.5B parameter model pre-trained on 800B tokens showed gains on math and coding when extra loops were run at inference, and Ouro carried the approach to 1.4B and 2.6B parameters on up to 7.7T tokens while pursuing adaptive computation, where harder inputs consume more depth.

The review's own open questions are about whether any of this is really thinking. A hidden state that changes across loops is not proof that one loop extracts, another reasons, and a third verifies. Until it is clear what each loop updates, whether loop count tracks problem difficulty, and whether a model can extrapolate past its training loop count, the outcome may hinge on whether stopping decisions can be learned stably — and for buyers comparing architectures, on whether the parameter savings survive contact with real inference costs.

FAQ
How does a Looped Transformer differ from an RNN?
Both reuse the same parameters to update a hidden state, but an RNN updates along the sequence direction. A Looped Transformer updates the same token sequence along the depth direction.
Does looping the same block reduce latency?
No. The review states parameter efficiency and compute efficiency are separate, and that sequential loops can run less efficiently than a normal deep Transformer on some hardware.
What was the first paper to use the name 'Looped Transformers'?
Giannou et al.'s 'Looped Transformers as Programmable Computers', posted to arXiv on January 30, 2023 and accepted at ICML 2023.
Qiita 機械学習Read Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleAI sped up creation, but the BPM still had to be felt by hand