AIToday
AI Business & IndustryAmazon AI BlogPublished: Sep 9, 2026, 04:00 JST2 min read

AWS G7 instances beat G5, G6 in LLM inference benchmarks

AWS G7 instances beat G5, G6 in LLM inference benchmarks

3 Key Points

  1. What happened

    AWS benchmarked two 30B Mixture-of-Experts models on SageMaker AI, comparing new G7 instances against older G5 and G6 families. G7, powered by NVIDIA Blackwell GPUs, won across throughput, latency, and cost tests.

  2. Why it matters

    In one test, ml.g7.12xlarge hit 391.3 output tokens per second, about 60.8% higher than G6 and 13.0% higher than G5, with lower latency. It did so using only two GPUs with 64 GB memory versus four GPUs with 96 GB on G5/G6.

  3. What to watch

    The gains stem partly from G7's native FP4 Tensor Core support for low-precision formats, which older instances lack. Whether these results hold for your workloads hinges on testing, since performance varies by model, token length, and concurrency.

WHO IT HITSEnterprise teams deploying LLMs on AWS SageMaker AI must re-evaluate their instance choices, as G7 offers clear price-performance gains over G5/G6 for real-time inference, though only in US East (Ohio) and US West (Oregon) currently.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

AWS's benchmark post targets a practical decision: which GPU instance to pick for LLM inference on SageMaker AI. With generative AI moving from experimentation to production, infrastructure costs directly impact the bottom line, making instance choice critical. The post's first test compares a coding model across G5, G6, and G7, while the second uses an automated recommendation tool to evaluate a different model across G6, G6e, and G7.

A key structural advantage is G7's native support for Blackwell's 4-bit floating-point format, which older generations run without hardware acceleration. G7 achieves its gains with fewer GPUs and less memory—two GPUs with 64 GB versus four with 96 GB on G5/G6—highlighting the generational jump in efficiency.

The stakes are clear for AWS customers: sticking with G5/G6 may mean paying more for less performance, but the gains are workload-dependent. G7 is currently limited to two AWS regions, and the best configuration varies by input length. The test results are a baseline; teams should benchmark their own workloads before committing.

FAQ
How much faster is G7 compared to G6?
In the coding model benchmark, ml.g7.12xlarge achieved about 60.8% higher throughput than G6 and 13.0% higher than G5. It also reduced average latency by about 37.6% compared with G6.
What is the cost savings with G7?
For the tested workloads, G7 provides approximately 1.4× lower cost per token than G6e, and 4.9× lower than G6 for chat workloads. The cheapest option was ml.g7.2xlarge at $0.90 per 1 million output tokens for chat.
Why is G7 so much better than older instances?
G7 instances have native FP4 Tensor Core support for the NVFP4 format, an advantage G5 and G6 lack. Only G7 instances support native FP4 Tensor Core support in the tested families.
Amazon AI BlogRead Original Article

Get the latest AI Business & Industry news every morning

For example, today's edition would include:

  • Jim Cramer picks Dell over Super Micro on accounting concernsYahoo Finance AI · 51m ago
  • Alphabet's AI spend sinks cash flow, but Cloud jumps 82%Yahoo Finance AI · 51m ago
  • Google's TimesFM-3 tops forecasting benchmarks, adds multivariate supportTHE DECODER · 51m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articlePerplexity CEO Aravind Srinivas: Banks want unpluggable AI