AIToday
Large Language ModelsAI Business & IndustryGIGAZINE AIPublished: Oct 2, 2026, 01:00 JST

Magnitude inference engine beats llama.cpp by up to 2x

Magnitude inference engine beats llama.cpp by up to 2x

3 Key Points

  1. What happened

    The Magnitude development team published an open-source inference engine that compiles and tunes model kernels on the PC itself, hitting about 2x llama.cpp's speed in certain conditions.

  2. Why it matters

    Running open models locally for AI agents has leaned on engines built for datacenter batching or broad hardware support, so Magnitude's on-device tuning looks aimed at a gap those tools leave.

  3. What to watch

    The headline 2x figure holds only in certain conditions, so the test is whether the gains repeat as the team adds Expert streaming and broader multi-device optimization.

WHO IT HITSDevelopers and small teams running AI agents on their own laptops or desktops are the most direct audience, since Magnitude targets exactly that setup. IT teams weighing local agent deployments may find the memory savings matter as much as raw speed.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The Magnitude team's starting point is a mismatch between how inference engines are usually built and how local AI agents actually run. Datacenter engines are optimized for batching large numbers of requests, while projects like llama.cpp prioritize covering as much hardware as possible. Neither, in the team's view, fits a PC running an agent for a long time — a workload that involves repeated file reads, command execution, and often several agents at once.

Their answer is to compile the small computational units called kernels on the machine that will run the model, and to tune parameters there too. The reasoning is that even within the same Apple silicon or NVIDIA family, performance and memory layouts differ by product, so device-specific tuning can beat a generic kernel built for wide compatibility. The team's benchmark used a 4bit-quantized Qwen 3.6 35B A3B at a 64K token context, comparing against llama.cpp across both a Mac and an NVIDIA DGX Spark.

Beyond raw speed, the release addresses long-running and parallel use. Magnitude allocates only the memory needed to hold the model's weights, grows that region as a session lengthens, and frees it when an agent stops, while a shared prefix cache reuses computation for system prompts common across sessions. The team has also flagged what comes next: Expert streaming, a compiler that decides kernel configuration automatically, and optimization across CPU, GPU, RAM, and storage together. Whether the early gains hold as those features land is the open question.

FAQ
How much faster is Magnitude than llama.cpp?
On an M4 Pro Mac with 48GB memory, generation speed was about 92% higher (57 tokens/second vs 30), while prefill rose about 9% (507 vs 466 tokens/second). On NVIDIA DGX Spark, generation was about 19% higher and prefill about 23% higher.
What hardware and agent tools does Magnitude support?
It runs on macOS, Windows, and Linux across Apple silicon, NVIDIA GPUs, AMD GPUs, and CPU-only machines. It connects to tools like Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline, plus an OpenAI-compatible API.
Can I use Magnitude offline?
Yes. Once a model is downloaded you can use it without an internet connection, and prompts and files stay on the PC. Magnitude itself is released under Apache License 2.0.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleDevOps engineer: Hermes Agent build cost him the thrill