AIToday
Large Language ModelsOpen-Source AIAI Business & Industryr/LocalLLaMAPublished: Mar 25, 2026, 16:57 JST1 min read

A single RTX 5090 GPU achieves impressive throughput speeds of 20 tokens/second for text generation and 700 tokens/second for prompt processing with the massive 397B-parameter Qwen3.5 model.

A single RTX 5090 GPU achieves impressive throughput speeds of 20 tokens/second for text generation and 700 tokens/second for prompt processing with the massive 397B-parameter Qwen3.5 model.

3 Key Points

  1. Qwen3.5-397B-A17B model in Q4_K_M quantization format runs on a single NVIDIA RTX 5090 with 256GB DDR4 RAM and AMD EPYC 7532 CPU

  2. Text generation speed reaches 20 tokens/second while prompt processing achieves 700 tokens/second throughput

  3. System uses PCIe 4.0 x16 connection, 2TB NVMe SSD storage, and llama-bench benchmarking tool with batch size of 8192

  4. Performance data addresses the lack of publicly available speed metrics for running 397B-parameter models on consumer/enthusiast single-GPU setups

Ask the AI about this article →

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Anthropic releases Claude Fable 5.1 and Mythos 5.1ITmedia AI+ · 1h ago
  • LLM serving: why continuous batching winsDaily Dose of Data Science · 1h ago
  • Anthropic's Claude Fable 5.1 Now on Snowflake Cortex AISnowflake AI Blog · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleVoxyflow introduces an AI bot capable of automating task execution directly from Kanban board cards