AIToday
Large Language ModelsOpen-Source AIQiita 機械学習Published: Oct 9, 2026, 13:00 JST

ax-engine hits 76 tok/s on M5 Max, beats mlx-lm

ax-engine hits 76 tok/s on M5 Max, beats mlx-lm

The open-source ax-engine for Apple Silicon Macs measured over 76 tok/s decode on one M5 Max setup, and about 2.4〜2.7x mlx-lm's decode speed in some conditions.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The project starts from a hardware quirk of Apple Silicon Macs: the GPU and CPU share the same Unified Memory Architecture. That means running a local LLM is often limited not by raw compute but by memory bandwidth, the rate at which the model's weights and KV cache can be read. In practice, a model may load fine yet generate tokens more slowly than expected.

ax-engine targets that gap by focusing on execution and serving rather than a chat UI or a cloud service. Its design bets include MTP decoding, in which an auxiliary prediction head proposes several tokens ahead and the main model verifies them, plus Prefix reuse and the ability to serve multiple models from a single process. The developer stresses that MTP's benefit is not constant — it shifts with the model architecture, whether MTP is supported, the quantization format, prompt and context length, the number of tokens generated, the Apple Silicon generation, and Unified Memory capacity. That is why the guidance is to compare like with like: same model, same quantization, same prompt.

The benchmark table itself is presented with placeholder values in this write-up, so the reproducible public figure is the reported M5 Max result of over 76 tok/s and the roughly 2.4 to 2.7x edge over mlx-lm under some conditions. The developer is explicit that these are not fixed numbers and can change widely by configuration. To widen the picture, they are asking for results from M1 through M5 machines with 16 GB up to 64 GB or more of memory, across quantization formats, along with chip and memory size, macOS and ax-engine versions, model name, context length, decode and prefill speeds, memory use, and any errors. Known limitations include uneven model support, MTP's dependence on model and settings, install and model-conversion issues, and no guarantee of compatibility with every OpenAI-compatible client.

FAQ
What exactly is ax-engine?
It is an open-source LLM inference and model serving engine for Apple Silicon Macs. It offers an OpenAI-compatible API, can serve multiple models from one process, and supports MTP and Prefix reuse on compatible models.
How much faster is it than mlx-lm?
In some conditions it measured roughly 2.4 to 2.7 times mlx-lm's decode speed, with over 76 tok/s on a specific M5 Max configuration. The developer says these are not fixed values and vary by model, quantization, context length, prompt and MTP settings.
Can I use my existing OpenAI-compatible tools with it?
Yes. Because ax-engine exposes an OpenAI-compatible API, any application that lets you set an OpenAI-compatible Base URL can point at the local ax-engine instead of a cloud API.
Qiita 機械学習Read Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleAnthropic bans sustained cruel abuse of Claude in policy update