AIToday
Large Language ModelsOpen-Source AIAI Business & IndustryFortune AIPublished: Aug 21, 2026, 19:03 JST3 min read

Chinese AI models seize OpenRouter majority as US firms defect on price

Chinese AI models seize OpenRouter majority as US firms defect on price

Key takeaway

  • Chinese AI models now dominate OpenRouter, a major routing platform, capturing over 60% of traffic in July.

  • American companies are switching to Chinese models for cost reasons, not coercion: DeepSeek costs one-twelfth what GPT-5.5 costs at comparable performance.

  • The article warns that most corporate AI strategies are stuck between expensive frontier models and cheap commodity alternatives, and that without open-weight competition from US labs, American firms risk losing the ecosystem to China.

3 Key Points

  1. What happened

    Chinese-developed models claimed all five top positions on OpenRouter in July, with Xiaomi's MiMo V2.5 ranked first by token volume. Chinese models now carry more than 60% of the platform's traffic, which exceeds 20 trillion tokens a week. By mid-July, Chinese models accounted for 58% of tokens processed by American firms on the platform—a shift from roughly 70% US model traffic a year ago.

  2. Why it matters

    US companies are choosing Chinese AI workload-by-workload because of cost, not compulsion. DeepSeek V4 Flash costs $0.14 per million input tokens versus $5.00 for GPT-5.5; Chinese open models run 60% to 90% cheaper than leading American offerings. Meanwhile, Alibaba's Qwen family has passed one billion cumulative downloads and replaced Meta's Llama as the most-downloaded open model family—a shift that locks developers into an ecosystem. Most Fortune 500 AI strategies are stuck in a "death zone": neither frontier-class nor cheap enough to compete.

  3. What to watch

    The article argues American labs must release frontier-class open-weight models on a regular schedule to compete. Without credible US open-weight alternatives, the risk is that "when the next generation of global software is built, [it will be] built on" Chinese models. Companies deploying hybrid routing—frontier models for high-stakes work, efficient open models for high-volume tasks—are cutting inference costs 60% to 90% on the majority of workloads.

Ask the AI about this article →

Context & Analysis

The article identifies a structural shift in AI adoption driven by economics rather than capability. While American labs (OpenAI, Anthropic, Google) still hold the frontier on the hardest reasoning tasks—GPT 5.5, Claude Fable 5, and Gemini 3.x lead on demanding enterprise work—the deployment market has bifurcated. Chinese labs, constrained by export controls on large GPU clusters, engineered around scarcity with token efficiency and novel architectures from the start. The cost gap is now decisive: American firms face a choice between paying premium prices for frontier capabilities they often do not need, or adopting cheaper alternatives. The author notes that Anthropic captures roughly half of total OpenRouter spending despite holding only 12% of token share—the "premium lane." Everything else is being crushed into either the commodity floor (cheap, open models) or disappearing from the middle.

The ecosystem lock-in effect mirrors historical platform shifts: developers optimize for what they can download and build tooling around what they deploy, just as Linux won servers and Android won phones. Alibaba's Qwen family passing one billion cumulative downloads signals that Chinese models are becoming the default foundation for developer projects globally. The article argues this outcome was not accidental: state support lowered effective costs, Xiaomi cut API prices by up to 99% in May, and the labs prioritized inference efficiency from day one. For American strategy, the implication is stark—the race is no longer just about frontier capability but about who owns the distribution and ecosystem.

FAQ

Which Chinese model ranked first on OpenRouter in July?
Xiaomi's MiMo V2.5 ranked first by token volume, followed by models from DeepSeek, MiniMax, Alibaba's Qwen family, and Moonshot's Kimi.
How much cheaper is DeepSeek V4 Flash than GPT-5.5?
DeepSeek V4 Flash costs $0.14 per million input tokens, compared with $5.00 for GPT-5.5. More broadly, Chinese open models run 60% to 90% cheaper than leading American offerings.
What happened to Meta's Llama model?
Alibaba's Qwen family replaced Meta's Llama as the most-downloaded open model family in the world. Llama, which defined open weight AI in 2023 and 2024, has fallen below 1% of routed volume on OpenRouter.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Nvidia revives Rubin CPX chip with major redesignYahoo Finance AI · 2h ago
  • AI advice followed by 79%, but well-being unchangedITmedia AI+ · 5h ago
  • Enterprises face agent governance gapSiliconANGLE AI · 8h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleApple cuts Vision team in VR retreat, pivots to AI glasses