AIToday
Large Language ModelsOpen-Source AILobsters AIPublished: Sep 29, 2026, 19:00 JST

MicroLLM lab lets users benchmark tiny 135M models in their browser

MicroLLM lab lets users benchmark tiny 135M models in their browser

3 Key Points

  1. What happened

    MicroLLM lab launched as a browser tool for benchmarking tiny LLMs, including 135M-parameter models, on objective checks, reporting speed in tokens/s and accuracy as a pass rate, with all numbers kept on the local machine.

  2. Why it matters

    The tool treats failures as the measurement itself — a 135M model is allowed to fail — and uses objective checks rather than writing-quality judgments, so results are meant to be comparable.

  3. What to watch

    Runs are confined to the user's browser and hardware, so scores are not directly comparable across devices. Watch which models and objective tests users actually select when the tool is used.

WHO IT HITSDevelopers and AI enthusiasts evaluating small models on their own laptops can now run objective benchmarks without sending data off-device. Because results are tied to each user's browser and hardware, comparisons across machines are limited.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

MicroLLM lab arrives as a browser-based alternative to benchmark suites that judge models on writing quality. Because its checks are objective — regex or exact tokens — a 135M model that fails a check is not a problem; the failure is the measurement the tool is designed to produce. The tool separates speed, reported as tokens/s of sustained decode plus suite wall time, from accuracy, reported as a pass rate on objective tests, and it uses the latest suite per model for its charts.

The setup also allows users to write their own benchmarks in JavaScript, with the editor eval()'d in the page origin so each check runs on the model's decoded text. There is an option to prompt a larger LLM, and an estimate uses the reader's last tok/s reading if one is available.

The main constraint is that results stay on the machine, which keeps data local but also means scores are shaped by each user's browser and hardware. How useful the tool proves will likely hinge on whether its objective checks become a common reference point, and on which models users choose to put through them.

FAQ
What kinds of checks does MicroLLM lab run?
It runs objective checks based on regex or exact tokens, not writing-quality judgments. A 135M model is allowed to fail — the failure is the measurement.
Where do the benchmark numbers go?
The numbers stay on your machine. Speed is measured in tokens/s sustained decode and suite wall time, and accuracy is a pass rate on objective tests.
How does the JavaScript benchmark editor work?
The editor is eval()'d in the page origin, and each check runs on the model's decoded text.

AI news that matters for your work, in one minute a day

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleOpenAI outlines "safety cases" for frontier AI training