
What happened
By September 2026, Amazon Bedrock, Azure AI Foundry and Vertex AI now resell frontier models at prices that diverge from $0.75/$3.75 per 1M tokens for Gemini 3.8 Flash to $10.00/$50.00 for GPT-6 Astra.
Why it matters
Buyers of on-demand model access face a wide gap in what they pay per token, so the same workload can cost far more on one platform than another.
What to watch
The low Flash-tier price is a fast, cheap tier, not a like-for-like reasoning competitor, so the real test is what Gemini 3.1 Pro costs once Google publishes a price.
WHO IT HITSCloud procurement and FinOps teams choosing where to run high-volume inference now face per-token bills that can differ by roughly an order of magnitude, so platform choice directly shapes operating cost.
Summaries like this, in your inbox every morning.
The three platforms started from the same pitch — pick a model, call an endpoint, pay per token — but the catalog counts show how differently they now count. Bedrock reports more than 100 models from 18-plus providers, Azure AI Foundry cites more than 10,000 with roughly 50 new additions a month, and Vertex AI's Model Garden sits past 200, with Azure's figure including many narrow open-weight variants and fine-tuned derivatives rather than just frontier releases.
The benchmark trackers tell a similarly split story. Anthropic's Claude Opus 5 tops SWE-bench Verified at 97.00%, yet Bedrock's own model cards still list Claude Opus 4.6 as the flagship Opus-tier offering, so Bedrock customers may be running a generation behind Anthropic's top-scoring release depending on the model ID they deploy. GPT-6 Astra leads on agent-oriented measures like Terminal-Bench v4.0, while Gemini 3.8 Flash is built for volume rather than the top of the leaderboard.
For a buyer, the decision hinges less on an abstract winner than on which constraint binds first. The headline price gap can be outweighed by Azure's regional multipliers, where an EU customer running GPT-6 Astra in the EU Data Zone now pays roughly 20% above global list price on top of a 9% increase applied earlier in the year. Capacity confirmation for launch-week access to models like GPT-6 Astra or Gemini 3.1 Pro looks likely to matter as much as list pricing.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Dell Technologies and Axelera AI announced a collaboration on the new Europa AI Processing Unit architecture f…

Dell took almost $61 billion in AI server orders last quarter and ended with a $95 billion backlog

In a post on X and other social platforms, Meta CEO Mark Zuckerberg said labs that fail to "focus on alignment…

Procter & Gamble is expanding its deployment of AI-based quality inspection technology in collaboration with S…

Apple finally built a smarter version of Siri, according to the WSJ

Bank of America CEO Brian Moynihan told investors that AI has helped the bank avoid layoffs
