AIToday
Large Language ModelsDaily Dose of Data SciencePublished: Oct 9, 2026, 10:01 JST

Uzu hits 92–117 tok/s on base M5, beating llama.cpp

Uzu hits 92–117 tok/s on base M5, beating llama.cpp

Uzu, Mirai's open-source engine, generated 92–117 tokens per second on a base M5 MacBook Pro with 4-bit Qwen3.5 9B, where llama.cpp hit 22 and MLX 25.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Local LLM inference runs at batch size one: one app, one active sequence, one device. That removes cross-user batching, but also narrows the problem — the engine optimizes a single decode on known hardware. Uzu attacks two quantities at once: bytes transferred per pass and accepted tokens per pass.

The bytes side starts with a hard memory-bandwidth limit. A base M5 moves 153 GB per second, and a 4-bit Qwen3.5 9B checkpoint weighs about 5.2 GB, capping ordinary decoding at roughly 153 ÷ 5.2 ≈ 29 tokens per second. Uzu's Medium checkpoint uses 4-bit asymmetric integer quantization, and its Metal kernels run the speculative verification step on the M5's int8 matrix path. A block-diagonal Random Hadamard Transform of 32 values spreads outliers before quantization, matching the Apple GPU kernel's SIMD width.

The tokens side uses DFlash and Weaver to draft 16-token trees, while rollback-free verification handles Qwen3.5's Gated DeltaNet layers — 24 of its 32 — by computing candidate outputs without touching live state, then committing only along the accepted path. Mirai's published 27B benchmark reports 2.1× MTPLX, 3.8× llama.cpp and 4.4× MLX, with speculation lifting 35.54 to 114 tokens per second. Speed still depends on the workload: code and maths accept more drafts than open-ended conversation.

FAQ
What makes Uzu faster than llama.cpp or MLX?
Uzu designs quantization, Metal kernels and speculative decoding as one system. It also uses DFlash and Weaver to propose long token sequences, then verifies them without repeatedly restoring model state.
Does Uzu work on older Macs?
Speculation is enabled per chip. The checkpoint's speculator configuration lists the M5, M5 Pro and M5 Max, so on M1 to M4 Macs the same command runs without a drafter and reports 1.00 t/f.
How do I install Uzu?
It installs through Homebrew and downloads the selected model on first use. Uzu can also expose an OpenAI-compatible HTTP endpoint on 127.0.0.1.
Daily Dose of Data ScienceRead Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleChatGPT's new dots agent runs 4.5-hour jobs while you sleep