
What happened
Open-weight models ran 56% of all tokens on Vercel AI Gateway in August, up from 7% in December, while OpenAI's Astra took 7.7% of all gateway spend in its first twelve days, more than twice Fable 5.1's 3.7%.
Why it matters
Open-weight models now carry more production traffic than all closed-weight models combined, so teams can reserve frontier models for tasks that justify the premium.
What to watch
Whether open-weight dollar share keeps climbing, since frontier models still kept the majority of spend in August. Watch the next monthly index for spend data.
WHO IT HITSEnterprise AI platform teams and engineering leads routing production workloads through Vercel AI Gateway are now able to run most of that traffic on open-weight models, while those paying for frontier models like Anthropic's Fable 5 are shifting workloads to cheaper tiers such as Opus 5.
Summaries like this, in your inbox every morning.
Vercel's AI Gateway Production Index tracks anonymized traffic routed through its gateway, which the company says routes tens of trillions of tokens monthly between production applications and AI labs. The September 2026 report covers data collected through August 2026, and includes updates logged between editions on September 17 and September 18. The core shift it documents is a pricing and workload migration: open-weight models, which processed fewer than one in ten tokens in December 2025, rose every month from April through August to reach 56% of volume. Over the same period, the average price per token fell 23.2% in August, the third straight monthly drop, and the median team running more than ten million tokens paid 7.6% less.
The report also traces how spending moved within Anthropic's lineup after Opus 5 came online at the end of July. Fable 5's share of gateway spend had surged to 13.2% in July when its access was restored under a lifted US export control, but fell to 4.9% in August as Opus 5's share rose to 22.5%. Nine in ten teams that ran Fable cut their usage, and Anthropic still kept 64 cents of every dollar spent in August because those workloads stepped down to Opus 5 rather than to another lab. A similar pattern played out at Google, where Gemini 3 Flash lost 95% of its share of gateway tokens since May and more than three-quarters of the volume it lost went to other labs.
The stakes here hinge on whether the migration toward open-weight and cheaper-tier models continues into September and beyond, and on whether frontier labs can keep premium spend even as their most expensive models lose share. For buyers running production AI workloads, the practical question is whether the gap between open-weight and frontier capability keeps narrowing enough to justify moving more tasks off the premium tiers.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Supply chain players say most AI data center cost models in China and the West count only electricity consumpt…

At IFA 2025 in Berlin, IFA Management CEO Leif Lindtner said Samsung Electronics left the fair after its leade…

Cocopelli introduced SAF, a generative-AI tool that searches internal rules, manuals and product materials and…

On September 17, 2026, KnowledgeSense added a function to ChatSense's Notebook that automatically gives AI-gen…

Broadcom CEO Hock Tan told Jim Cramer on Mad Money that demand for AI compute infrastructure for development a…

Netskope shares rose nearly 19% this week, helped by an overweight rating and $21 per share target from Stephe…
