AIToday
Large Language ModelsTop Companies' AI MovesAI Business & IndustryTop Companies AIPublished: Aug 24, 2026, 06:32 JST2 min read

RTX 5090 Laptop GPU Beats M5 Max in LLM Tests

RTX 5090 Laptop GPU Beats M5 Max in LLM Tests

Key takeaway

  • NVIDIA's laptop RTX 5090 beats Apple's M5 Max in LLM tests for smaller models.

  • It achieves up to 133% faster token generation.

  • But with larger models or longer context, the M5 Max wins.

3 Key Points

  1. What happened

    In benchmark tests by Alex Ziskind, NVIDIA's laptop RTX 5090 outperformed Apple's M5 Max in prompt processing and token generation for several LLMs, achieving up to 133% faster performance in some workloads.

  2. Why it matters

    The RTX 5090's 24GB VRAM limits its performance with larger models or longer context lengths. When context length increased, the M5 Max's unified memory (128GB) gave it a clear lead, such as 181.40 tokens/second vs 25.28 for the RTX 5090.

  3. What to watch

    The RTX 5090 excels with smaller models like Gemma 4 12B, Qwen3.5 27B, and Qwen3.6 35B A3B, but for dense models like Qwen3.5 122B, it is not viable. NVIDIA's RTX Spark launch may target heavier AI models.

Ask the AI about this article →

Context & Analysis

The benchmark results highlight a key difference in architecture between discrete GPUs and Apple's unified memory. The RTX 5090's raw compute power shines when the model fits within its 24GB VRAM, delivering up to 133% faster Performance in token generation for smaller LLMs. However, as context length grows, the VRAM fills up, tanking performance to 25.28 tokens/second, while the M5 Max's 128GB unified memory keeps it at 181.40 tokens/second.

This trade-off is central to choosing hardware for AI workloads. For users running large models or long contexts, the M5 Max is the practical choice. But for typical local LLM tasks with smaller models, the RTX 5090 offers superior speed, which may justify its cost for enthusiasts. NVIDIA's upcoming RTX Spark, which the article suggests is partly motivated by running heavier AI models, could address this gap, but details are not provided.

FAQ

What specific tests did the RTX 5090 vs M5 Max run?
They ran Qwen3 4B with 4-bit quantization, and also Gemma 4 12B, Qwen3.5 27B, and Qwen3.6 35B A3B, plus Qwen3.5 122B.
Why does the M5 Max outperform the RTX 5090 in some tests?
The M5 Max has 128GB of unified memory, which handles larger models and longer context lengths better, while the RTX 5090's 24GB VRAM runs out.
Top Companies AIRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • DataAgent launches with $10M to auto-fix Kubernetes faultsSiliconANGLE AI · 2h ago
  • SK Hynix custom HBM boosts inference up to 5.15xDIGITIMES Asia · 2h ago
  • Nvidia Earnings: Boring by Design, Avoiding a Consolidated WorldStratechery (Ben Thompson) · 2h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAjinomoto Uses AI for PR, Sparks Controversy