AIToday
Large Language ModelsZenn AI/MLPublished: Sep 28, 2026, 22:00 JST

Claude Fable 5.1 builds matrix-free Transformer site, then a 10M Japanese SLM

Claude Fable 5.1 builds matrix-free Transformer site, then a 10M Japanese SLM

3 Key Points

  1. What happened

    Mizushima (technology evangelist at Nextbeat) gave Claude Fable 5.1 a five-step goal chain; it first shipped a 14-chapter "Transformers without matrices" tutorial on GitHub Pages, then a separate repo with a 1M-parameter Japanese SLM written in Scala 3 with hand-written backprop.

  2. Why it matters

    The 1M model trained on 313 Aozora Bunko works (about 6.41 million tokens, roughly 1.9 epochs) produced grammatically shaped Japanese with correct particles and punctuation, suggesting matrix-free, derivative-free training can reach sentence-level language modeling.

  3. What to watch

    The 10M model reached validation loss 2.70, but the article says CPU training is "not a realistic approach" for large-scale learning — the test is whether the same recipe scales efficiently beyond this size.

WHO IT HITSFor engineers and educators exploring low-dependency, CPU-only language-model training, this shows a viable path — but one whose speed (a claimed 19× improvement to 27,700 tokens/sec at 1M parameters) still leaves it unfit for large-scale work. ML researchers may want to probe the accuracy of the article's conceptual claims.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The site and the SLMs emerged from a single goal-setting workflow. The author, whose graduate work was in parsing and programming-language processing, first asked for a tutorial, then asked whether derivatives could be dropped from the explanation, then asked for a 1M-parameter SLM, then for speeding up, then for scale — each step building on the previous output.

A notable intermediate finding shaped the rest: on the same roughly 7,000-parameter model and corpus, perturbation-based (zeroth-order) training reached loss 1.0 in 62 minutes versus 4 minutes for backprop with Adam — about a 15× gap. That result drove the site's switch to a "no derivatives" framing, with backprop relegated to an appendix. Speed work came before any matrix library: loop reordering lifted 16-thread throughput from 1,430 to 5,350 tokens/sec, then SIMD and Float32 conversion pushed it to 27,700. The author also cautions that Zen 4's AVX-512 splits 512-bit registers across 256-bit units, so the SIMD gain came from fewer instructions, not wider arithmetic.

What the exercise ultimately tests is whether the pedagogy holds up as engineering — whether explaining Transformers without matrices, and training small models without matrix libraries, can still produce recognizable language output. The evidence is mixed but concrete: the 1M model's sentences have structure without meaning, while the 10M model stabilizes speaker-verb agreement and reuses character names within a scene. Whether that arc continues to larger sizes on CPU, or requires the very matrix infrastructure the project set out to avoid, is the open question the author leaves to readers.

FAQ
How fast did training get after optimization?
At 1M parameters, the model went from 1,430 tokens/sec (16 threads, first version) to 27,700 tokens/sec after loop reordering, SIMD, and Float32 conversion — a 19× speedup, with no matrix library used at any point.
What data was the 10M model trained on?
The author removed the per-author limit and pulled 3,911 Aozora Bunko works (about 71 million characters). Training ran 4,000 steps (about 33 million tokens, 0.5 epochs) in 2 hours 50 minutes, reaching validation loss 2.70.
Does the article claim the explanations are fully verified?
No. The author notes his background is not in ML, says the evaluation was limited to checking whether generated text looked reasonable, and invites experts to point out any errors.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Autoheal raises $7.9 million for self-fixing AI agentsSiliconANGLE AI · 43m ago
  • Paul Cheek: 30% of S&P 500 execs AI-literate, 78% gapFortune AI · 43m ago
  • Agent cost per successful task: a Zenn design-variable argumentZenn AI/ML · 43m ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleOpenAI Knew Pirated Books Risked Authors' Livelihoods