
What happened
Mizushima (technology evangelist at Nextbeat) gave Claude Fable 5.1 a five-step goal chain; it first shipped a 14-chapter "Transformers without matrices" tutorial on GitHub Pages, then a separate repo with a 1M-parameter Japanese SLM written in Scala 3 with hand-written backprop.
Why it matters
The 1M model trained on 313 Aozora Bunko works (about 6.41 million tokens, roughly 1.9 epochs) produced grammatically shaped Japanese with correct particles and punctuation, suggesting matrix-free, derivative-free training can reach sentence-level language modeling.
What to watch
The 10M model reached validation loss 2.70, but the article says CPU training is "not a realistic approach" for large-scale learning — the test is whether the same recipe scales efficiently beyond this size.
WHO IT HITSFor engineers and educators exploring low-dependency, CPU-only language-model training, this shows a viable path — but one whose speed (a claimed 19× improvement to 27,700 tokens/sec at 1M parameters) still leaves it unfit for large-scale work. ML researchers may want to probe the accuracy of the article's conceptual claims.
Summaries like this, in your inbox every morning.
The site and the SLMs emerged from a single goal-setting workflow. The author, whose graduate work was in parsing and programming-language processing, first asked for a tutorial, then asked whether derivatives could be dropped from the explanation, then asked for a 1M-parameter SLM, then for speeding up, then for scale — each step building on the previous output.
A notable intermediate finding shaped the rest: on the same roughly 7,000-parameter model and corpus, perturbation-based (zeroth-order) training reached loss 1.0 in 62 minutes versus 4 minutes for backprop with Adam — about a 15× gap. That result drove the site's switch to a "no derivatives" framing, with backprop relegated to an appendix. Speed work came before any matrix library: loop reordering lifted 16-thread throughput from 1,430 to 5,350 tokens/sec, then SIMD and Float32 conversion pushed it to 27,700. The author also cautions that Zen 4's AVX-512 splits 512-bit registers across 256-bit units, so the SIMD gain came from fewer instructions, not wider arithmetic.
What the exercise ultimately tests is whether the pedagogy holds up as engineering — whether explaining Transformers without matrices, and training small models without matrix libraries, can still produce recognizable language output. The evidence is mixed but concrete: the 1M model's sentences have structure without meaning, while the 10M model stabilizes speaker-verb agreement and reuses character names within a scene. Whether that arc continues to larger sizes on CPU, or requires the very matrix infrastructure the project set out to avoid, is the open question the author leaves to readers.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Autoheal AI Inc. raised $7.9 million in seed funding led by Innovation Endeavors, with Emergent Ventures, U&I…
Paul Cheek's AI-Driven Enterprise Institute study found just over 30% of S&P 500 executives are AI-literate, a…

From 7/30 to 9/17, /code-review ran 23 times with at most 1 subagent; from 9/23 it launched 10 at once, hittin…

A Zenn article narrowed agent cost design to three topics: cache depends on prefix stability, routing should b…

Working alone with 10 parallel Claude Code sessions, he logged 2,848 commits, 1,212 pull requests and 1,138 me…

Two Claude Code scheduled tasks on 9:10 and 10:01 morning runs produced no start rows, no errors and no notifi…
