AIToday
Large Language ModelsAI Coding AssistantsZenn AI/MLPublished: Oct 2, 2026, 22:00 JST

Tiny LLM hits 573,952 params, repeat-mumbles at 18000字

Tiny LLM hits 573,952 params, repeat-mumbles at 18000字

3 Key Points

  1. What happened

    A hobbyist built a fully from-scratch tiny language model with 573,952 parameters, a 16-token context, and byte-level tokenization, trained on his own unpublished novel on a single CPU thread.

  2. Why it matters

    At roughly 300 training runs it memorized a short greeting, but at 18,000字 the model did not learn grammar, instead repeating one character, suggesting the tiny context never captured prior text.

  3. What to watch

    The author blames the short context and thin evaluation for the repetition, and says a model that tracks what it has already said could avoid the loop; he plans a next version.

WHO IT HITSThis is a hobbyist learning project rather than a product, so it lands mainly on developers and students who want a hands-on feel for how language models are trained, not on teams buying AI tools.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The project was framed from the start as a record of an experiment rather than a tutorial, and the author set three hard constraints: no external code libraries, a single CPU thread, and everything self-contained. The stated reason for the CPU limit was that parallel processing would blur the study of how an LLM works, while the from-scratch rule reflected both curiosity and the fact that no existing library targets such a stripped-down setup.

The model itself is deliberately simple, with three layers, sigmoid activation, and argmax token selection, trained by backpropagation and gradient descent with mean squared error. The early results show memorization working on short inputs, but as the training text grew the output degraded into repeated characters, which the author connects to the 16-token context and a scoring method that does not discourage repetition.

The stakes here are personal rather than commercial. The author describes the work as about two days driven mostly by curiosity, and the failures he lists, such as weight matrices initialized to 1.0 and a context too short to see enough Japanese characters, are the kind of practical lessons that may carry into the next model he says he intends to build.

FAQ
How big is the model?
It has 573,952 total parameters, which the author describes as a few millionths of the latest LLMs, with a context length of 16 and byte-level tokenization.
What happened when he trained it on his novel?
With a roughly 180-character passage trained about 300 times, the model broke down; even at around 1,800 characters after three runs it still repeated garbled phrases, and at 18,000 characters it did not improve, so he stopped.
Why does the model repeat itself?
The author points to the very short context length, which shows only a few preceding Japanese characters, and to an evaluation method that gives no penalty for repeating the same phrase.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleHao Global admits $160 million GPU smuggling to China