
NVIDIA's laptop RTX 5090 beats Apple's M5 Max in LLM tests for smaller models.
It achieves up to 133% faster token generation.
But with larger models or longer context, the M5 Max wins.
What happened
In benchmark tests by Alex Ziskind, NVIDIA's laptop RTX 5090 outperformed Apple's M5 Max in prompt processing and token generation for several LLMs, achieving up to 133% faster performance in some workloads.
Why it matters
The RTX 5090's 24GB VRAM limits its performance with larger models or longer context lengths. When context length increased, the M5 Max's unified memory (128GB) gave it a clear lead, such as 181.40 tokens/second vs 25.28 for the RTX 5090.
What to watch
The RTX 5090 excels with smaller models like Gemma 4 12B, Qwen3.5 27B, and Qwen3.6 35B A3B, but for dense models like Qwen3.5 122B, it is not viable. NVIDIA's RTX Spark launch may target heavier AI models.
Ask the AI about this article →
The benchmark results highlight a key difference in architecture between discrete GPUs and Apple's unified memory. The RTX 5090's raw compute power shines when the model fits within its 24GB VRAM, delivering up to 133% faster Performance in token generation for smaller LLMs. However, as context length grows, the VRAM fills up, tanking performance to 25.28 tokens/second, while the M5 Max's 128GB unified memory keeps it at 181.40 tokens/second.
This trade-off is central to choosing hardware for AI workloads. For users running large models or long contexts, the M5 Max is the practical choice. But for typical local LLM tasks with smaller models, the RTX 5090 offers superior speed, which may justify its cost for enthusiasts. NVIDIA's upcoming RTX Spark, which the article suggests is partly motivated by running heavier AI models, could address this gap, but details are not provided.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Israeli startup DataAgent Ltd
SK Hynix presented a custom HBM concept at SEMICON Taiwan 2026, where compute functions are placed in the base…

Taoyuan is positioning itself as a northern hub for AI data centers (AIDC), citing the Tatan area and an LNG c…

The U.S. Department of Defense announced on August 31 that it has deployed ChatGPT Mil, a customized version o…

Nvidia reported earnings that were both remarkable and boring, reflecting its focus on avoiding a consolidated…

Anthropic has agreed to a $35bn cloud-computing contract with Lambda, a Nvidia-backed cloud provider
