
What happened
AWS benchmarked two 30B Mixture-of-Experts models on SageMaker AI, comparing new G7 instances against older G5 and G6 families. G7, powered by NVIDIA Blackwell GPUs, won across throughput, latency, and cost tests.
Why it matters
In one test, ml.g7.12xlarge hit 391.3 output tokens per second, about 60.8% higher than G6 and 13.0% higher than G5, with lower latency. It did so using only two GPUs with 64 GB memory versus four GPUs with 96 GB on G5/G6.
What to watch
The gains stem partly from G7's native FP4 Tensor Core support for low-precision formats, which older instances lack. Whether these results hold for your workloads hinges on testing, since performance varies by model, token length, and concurrency.
WHO IT HITSEnterprise teams deploying LLMs on AWS SageMaker AI must re-evaluate their instance choices, as G7 offers clear price-performance gains over G5/G6 for real-time inference, though only in US East (Ohio) and US West (Oregon) currently.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
AWS's benchmark post targets a practical decision: which GPU instance to pick for LLM inference on SageMaker AI. With generative AI moving from experimentation to production, infrastructure costs directly impact the bottom line, making instance choice critical. The post's first test compares a coding model across G5, G6, and G7, while the second uses an automated recommendation tool to evaluate a different model across G6, G6e, and G7.
A key structural advantage is G7's native support for Blackwell's 4-bit floating-point format, which older generations run without hardware acceleration. G7 achieves its gains with fewer GPUs and less memory—two GPUs with 64 GB versus four with 96 GB on G5/G6—highlighting the generational jump in efficiency.
The stakes are clear for AWS customers: sticking with G5/G6 may mean paying more for less performance, but the gains are workload-dependent. G7 is currently limited to two AWS regions, and the best configuration varies by input length. The test results are a baseline; teams should benchmark their own workloads before committing.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Alphabet's second-quarter revenue grew 24% to $119.8 billion, while free cash flow turned negative $5.9 billio…

On September 8, Mad Money's Jim Cramer said Super Micro has accounting "irregularities," said he cannot recomm…

Google Research released TimesFM-3, a 330 million-parameter forecasting model trained on over one trillion dat…

Twenty-five Fields Medal winners, including Terence Tao, signed a joint statement warning that AI companies tr…

Meta is asking individual contributors in its Applied AI division whether they want to return to manager roles…

Todd Hughes, who trains language tutors at Rosetta Stone, told Fortune that AI can build vocabulary and aid co…
