AIToday
Large Language ModelsOpen-Source AIAI Business & IndustryVercel AI BlogPublished: Sep 19, 2026, 10:00 JST

Open-weight models hit 56% of AI Gateway tokens

Open-weight models hit 56% of AI Gateway tokens

3 Key Points

  1. What happened

    Open-weight models ran 56% of all tokens on Vercel AI Gateway in August, up from 7% in December, while OpenAI's Astra took 7.7% of all gateway spend in its first twelve days, more than twice Fable 5.1's 3.7%.

  2. Why it matters

    Open-weight models now carry more production traffic than all closed-weight models combined, so teams can reserve frontier models for tasks that justify the premium.

  3. What to watch

    Whether open-weight dollar share keeps climbing, since frontier models still kept the majority of spend in August. Watch the next monthly index for spend data.

WHO IT HITSEnterprise AI platform teams and engineering leads routing production workloads through Vercel AI Gateway are now able to run most of that traffic on open-weight models, while those paying for frontier models like Anthropic's Fable 5 are shifting workloads to cheaper tiers such as Opus 5.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Context & Analysis

Vercel's AI Gateway Production Index tracks anonymized traffic routed through its gateway, which the company says routes tens of trillions of tokens monthly between production applications and AI labs. The September 2026 report covers data collected through August 2026, and includes updates logged between editions on September 17 and September 18. The core shift it documents is a pricing and workload migration: open-weight models, which processed fewer than one in ten tokens in December 2025, rose every month from April through August to reach 56% of volume. Over the same period, the average price per token fell 23.2% in August, the third straight monthly drop, and the median team running more than ten million tokens paid 7.6% less.

The report also traces how spending moved within Anthropic's lineup after Opus 5 came online at the end of July. Fable 5's share of gateway spend had surged to 13.2% in July when its access was restored under a lifted US export control, but fell to 4.9% in August as Opus 5's share rose to 22.5%. Nine in ten teams that ran Fable cut their usage, and Anthropic still kept 64 cents of every dollar spent in August because those workloads stepped down to Opus 5 rather than to another lab. A similar pattern played out at Google, where Gemini 3 Flash lost 95% of its share of gateway tokens since May and more than three-quarters of the volume it lost went to other labs.

The stakes here hinge on whether the migration toward open-weight and cheaper-tier models continues into September and beyond, and on whether frontier labs can keep premium spend even as their most expensive models lose share. For buyers running production AI workloads, the practical question is whether the gap between open-weight and frontier capability keeps narrowing enough to justify moving more tasks off the premium tiers.

FAQ
What happened to Fable 5's share after Opus 5 arrived?
Fable 5 fell from 13.2% of gateway spend in July to 4.9% in August as Opus 5's share rose to 22.5%. More teams moved Fable workloads to Opus 5 than to any other model.
How many teams used Jev in its first day?
Within its first 24 hours, Jev was being used by nearly 13% of paid teams. That was 2x as many as the GPT-5.6 family and over 6x as many as Fable 5.1.
Vercel AI BlogRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Cocopelli deploys generative AI tool SAF at Toyokawa Shinkin BankITmedia AI+ · 1h ago
  • KnowledgeSense adds auto-naming to ChatSense NotebookITmedia AI+ · 1h ago
  • TypeSafe AI's Jev cracks the LLM monolithHacker News · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleOpenAI, Microsoft docs warned of 'doom loop' for web