AIToday
Large Language ModelsTop Companies' AI MovesAI Business & IndustryTop Companies AIPublished: Sep 17, 2026, 06:30 JST

Amazon Bedrock, Azure AI Foundry, Vertex AI diverge on price

Amazon Bedrock, Azure AI Foundry, Vertex AI diverge on price

3 Key Points

  1. What happened

    By September 2026, Amazon Bedrock, Azure AI Foundry and Vertex AI now resell frontier models at prices that diverge from $0.75/$3.75 per 1M tokens for Gemini 3.8 Flash to $10.00/$50.00 for GPT-6 Astra.

  2. Why it matters

    Buyers of on-demand model access face a wide gap in what they pay per token, so the same workload can cost far more on one platform than another.

  3. What to watch

    The low Flash-tier price is a fast, cheap tier, not a like-for-like reasoning competitor, so the real test is what Gemini 3.1 Pro costs once Google publishes a price.

WHO IT HITSCloud procurement and FinOps teams choosing where to run high-volume inference now face per-token bills that can differ by roughly an order of magnitude, so platform choice directly shapes operating cost.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Context & Analysis

The three platforms started from the same pitch — pick a model, call an endpoint, pay per token — but the catalog counts show how differently they now count. Bedrock reports more than 100 models from 18-plus providers, Azure AI Foundry cites more than 10,000 with roughly 50 new additions a month, and Vertex AI's Model Garden sits past 200, with Azure's figure including many narrow open-weight variants and fine-tuned derivatives rather than just frontier releases.

The benchmark trackers tell a similarly split story. Anthropic's Claude Opus 5 tops SWE-bench Verified at 97.00%, yet Bedrock's own model cards still list Claude Opus 4.6 as the flagship Opus-tier offering, so Bedrock customers may be running a generation behind Anthropic's top-scoring release depending on the model ID they deploy. GPT-6 Astra leads on agent-oriented measures like Terminal-Bench v4.0, while Gemini 3.8 Flash is built for volume rather than the top of the leaderboard.

For a buyer, the decision hinges less on an abstract winner than on which constraint binds first. The headline price gap can be outweighed by Azure's regional multipliers, where an EU customer running GPT-6 Astra in the EU Data Zone now pays roughly 20% above global list price on top of a 9% increase applied earlier in the year. Capacity confirmation for launch-week access to models like GPT-6 Astra or Gemini 3.1 Pro looks likely to matter as much as list pricing.

FAQ
How much does the same daily workload cost on each platform?
A workload of 10 million input and 2 million output tokens a day costs roughly $100 a day on Claude Opus 4.6 through Bedrock, about $200 a day on GPT-6 Astra's short-context tier, and about $15 a day on Gemini 3.8 Flash.
Which platform has the largest model catalog?
Microsoft describes Azure AI Foundry as having more than 10,000 models with roughly 50 new additions a month, while Bedrock lists more than 100 models from 18-plus providers and Vertex AI's Model Garden is past 200.
Is the cheapest flagship model also the best on benchmarks?
No. Gemini 3.8 Flash scores 80.0% on SWE-bench Verified, while Claude Opus 5 tops that leaderboard at 97.00%; Flash is designed as a cheap, fast, high-volume tier rather than a reasoning competitor.
Top Companies AIRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Intel Xeon gains 2.4x in MLPerf v6.1 via softwareTop Companies AI · 3h ago
  • Apple's Smarter Siri Still Lags in AI Race, WSJTop Companies AI · 3h ago
  • Mark Zuckerberg: labs ignoring "focus on alignment will fall behind"Top Companies AI · 3h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articlerenue, Marubeni AI app cuts blueprint quantity takeoff from 12 hours to 10 minutes