
What happened
MicroLLM lab launched as a browser tool for benchmarking tiny LLMs, including 135M-parameter models, on objective checks, reporting speed in tokens/s and accuracy as a pass rate, with all numbers kept on the local machine.
Why it matters
The tool treats failures as the measurement itself — a 135M model is allowed to fail — and uses objective checks rather than writing-quality judgments, so results are meant to be comparable.
What to watch
Runs are confined to the user's browser and hardware, so scores are not directly comparable across devices. Watch which models and objective tests users actually select when the tool is used.
WHO IT HITSDevelopers and AI enthusiasts evaluating small models on their own laptops can now run objective benchmarks without sending data off-device. Because results are tied to each user's browser and hardware, comparisons across machines are limited.
Summaries like this, in your inbox every morning.
MicroLLM lab arrives as a browser-based alternative to benchmark suites that judge models on writing quality. Because its checks are objective — regex or exact tokens — a 135M model that fails a check is not a problem; the failure is the measurement the tool is designed to produce. The tool separates speed, reported as tokens/s of sustained decode plus suite wall time, from accuracy, reported as a pass rate on objective tests, and it uses the latest suite per model for its charts.
The setup also allows users to write their own benchmarks in JavaScript, with the editor eval()'d in the page origin so each check runs on the model's decoded text. There is an option to prompt a larger LLM, and an estimate uses the reader's last tok/s reading if one is available.
The main constraint is that results stay on the machine, which keeps data local but also means scores are shaped by each user's browser and hardware. How useful the tool proves will likely hinge on whether its objective checks become a common reference point, and on which models users choose to put through them.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Stratechery published a piece after OpenAI's Dev Day, saying the showcased product is, frankly, pretty confusi…

Electronic Library and NTT Docomo Business co-developed "ELNET AI" and will launch the full version on October…

KnowledgeSense said on September 30 that CodeSense, a Japanese-made AI agent, will ship within a few weeks, an…

Instinct, founded last year by 23-year-old Noah Shinn, raised $1 billion in a Series C, lifting its valuation…

Palo Alto Networks CEO Nikesh Arora let an unreleased Anthropic model, Mythos, probe his company's internal sy…

SAIL generates a trajectory with a VLM, runs it in a simulator, scores progress from video, and sends the corr…
