
The open-source ax-engine for Apple Silicon Macs measured over 76 tok/s decode on one M5 Max setup, and about 2.4〜2.7x mlx-lm's decode speed in some conditions.
Summaries like this, in your inbox every morning.
The project starts from a hardware quirk of Apple Silicon Macs: the GPU and CPU share the same Unified Memory Architecture. That means running a local LLM is often limited not by raw compute but by memory bandwidth, the rate at which the model's weights and KV cache can be read. In practice, a model may load fine yet generate tokens more slowly than expected.
ax-engine targets that gap by focusing on execution and serving rather than a chat UI or a cloud service. Its design bets include MTP decoding, in which an auxiliary prediction head proposes several tokens ahead and the main model verifies them, plus Prefix reuse and the ability to serve multiple models from a single process. The developer stresses that MTP's benefit is not constant — it shifts with the model architecture, whether MTP is supported, the quantization format, prompt and context length, the number of tokens generated, the Apple Silicon generation, and Unified Memory capacity. That is why the guidance is to compare like with like: same model, same quantization, same prompt.
The benchmark table itself is presented with placeholder values in this write-up, so the reproducible public figure is the reported M5 Max result of over 76 tok/s and the roughly 2.4 to 2.7x edge over mlx-lm under some conditions. The developer is explicit that these are not fixed numbers and can change widely by configuration. To widen the picture, they are asking for results from M1 through M5 machines with 16 GB up to 64 GB or more of memory, across quantization formats, along with chip and memory size, macOS and ax-engine versions, model name, context length, decode and prefill speeds, memory use, and any errors. Known limitations include uneven model support, MTP's dependence on model and settings, install and model-conversion issues, and no guarantee of compatibility with every OpenAI-compatible client.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Stanford and Carnegie Mellon researchers tracked 1182 Character.AI users from September 2024 to August 2025, t…

A study presented at USENIX Security Symposium 2026 examined 196,682 resumes from hiring platforms and detecte…

Google Cloud announced Gemini Agent, a cloud-based multi-agent tool

Google Cloud announced Gemini agent, a universal agent that handles answers, knowledge work, media creation, a…

The guide shows code that pulls AAPL prices with yfinance, averages tweet sentiment via TextBlob, then fits a…

Anthropic's October 8, 2026 update adds a ban on sustained and unnecessary abusive or cruel behavior toward Cl…
