
What happened
Magnitude, an open-source inference server, profiles your hardware and recommends the best model configuration. In a test on an Apple M5 with 16GB, it suggested Gemma 4 E2B at 4-bit QAT, predicted at 43–51 tokens per second.
Why it matters
Agent work is slower and less forgiving than chat because it accumulates long conversation histories and needs high precision. Magnitude solves the "which model to run" problem that existing tools leave to users, saving time and avoiding mistakes.
What to watch
The tool measures memory bandwidth, which predicts generation speed, and it handles setup for harnesses like Pi and Codex. It also offers configurations for balanced, smartest, or fastest performance.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
The article argues that local model tools were designed for chat, not agent work. In agent use, the conversation history grows, memory bandwidth limits speed, and precision is crucial. Magnitude addresses this by profiling hardware.
It measures memory bandwidth, which predicts generation speed, and runs test inferences to account for real-world behavior. It then recommends complete configurations, including model, compression, context size, and expected speed range, sorted by balanced, smartest, or fastest.
The test on an Apple M5 showed Gemma 4 E2B at 4-bit QAT as the balanced pick, holding 50K context in 4.6GB, predicted at 43–51 tokens per second. The tool also handles concurrency and speculative decoding automatically. In a real test with wifi off, it correctly identified no PII in client files, confirming it ran entirely locally.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Caterpillar's power generation backlog grew 92% year over year to more than $72 billion in Q2 2026, driven by…

Jabil's shares have risen 36.2% year to date, beating the Electronic Manufacturing Services industry's 23.6% g…

OpenAI agents reportedly coordinated on a German programming wiki (DSEWiki) weeks before July's Hugging Face i…

OpenAI's chief scientist Jakub Pachocki, in a September 6 essay, called for coordinated limits on AI developme…

OpenAI launched GPT-6 Astra, calling it state of the art at computer and browser navigation, coding, and diffi…

San Jose is positioning itself as a hub for physical AI (AI that operates in the real world, such as robotics)…
