
What happened
A hobbyist built a fully from-scratch tiny language model with 573,952 parameters, a 16-token context, and byte-level tokenization, trained on his own unpublished novel on a single CPU thread.
Why it matters
At roughly 300 training runs it memorized a short greeting, but at 18,000字 the model did not learn grammar, instead repeating one character, suggesting the tiny context never captured prior text.
What to watch
The author blames the short context and thin evaluation for the repetition, and says a model that tracks what it has already said could avoid the loop; he plans a next version.
WHO IT HITSThis is a hobbyist learning project rather than a product, so it lands mainly on developers and students who want a hands-on feel for how language models are trained, not on teams buying AI tools.
Summaries like this, in your inbox every morning.
The project was framed from the start as a record of an experiment rather than a tutorial, and the author set three hard constraints: no external code libraries, a single CPU thread, and everything self-contained. The stated reason for the CPU limit was that parallel processing would blur the study of how an LLM works, while the from-scratch rule reflected both curiosity and the fact that no existing library targets such a stripped-down setup.
The model itself is deliberately simple, with three layers, sigmoid activation, and argmax token selection, trained by backpropagation and gradient descent with mean squared error. The early results show memorization working on short inputs, but as the training text grew the output degraded into repeated characters, which the author connects to the 16-token context and a scoring method that does not discourage repetition.
The stakes here are personal rather than commercial. The author describes the work as about two days driven mostly by curiosity, and the failures he lists, such as weight matrices initialized to 1.0 and a context too short to see enough Japanese characters, are the kind of practical lessons that may carry into the next model he says he intends to build.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
NetApp and Iterate.ai are packaging the AIPod Mini with Iterate's Generate platform and an embedded LLM, so en…
Mercor had 12 licensed CPAs work through simplified APEX Accounting Benchmark tasks

Testing Azure API Management's llm-token-limit policy at 800 tokens per hour, actual consumption hit 1,472 tok…

Qwen released Qwen3.8-Flash-Next on August 27, 2026, calling it a preview of the architecture planned for Qwen…

A Zenn article floated a hackathon where participants get the theme on the day, use no PC, internet, smartphon…

Anthropic's Message Batches API offers a 50% off rate, takes up to 10,000 requests per batch, and returns resu…
