AIToday
Large Language ModelsTHE DECODERPublished: Aug 21, 2026, 01:00 JST4 min read

Chinese AI models now match Western leaders on most benchmarks

Chinese AI models now match Western leaders on most benchmarks

Key takeaway

  • Chinese AI labs have closed the performance gap with Western leaders across most benchmarks—Moonshot's K3 and Alibaba's Qwen3.8-Max now rank near the top on broad evaluations of reasoning, coding, and long-context tasks.

  • The Western lead survives only in narrow areas: abstract pattern recognition (where margins vary by test), agentic reliability (Opus 5 at 54% vs. K3 at 39% on repeated-run tasks), and offensive cybersecurity (though that gap is narrowing fast).

  • Investors worry that without exclusive capabilities, model performance alone cannot sustain a business advantage.

3 Key Points

  1. What happened

    Chinese models—Moonshot's Kimi K3, Alibaba's Qwen3.8-Max, and GLM-5.3—have closed the performance gap with OpenAI and Anthropic. K3 ranks third on Artificial Analysis's Intelligence Index with 57 points (behind Opus 5 at 61 and GPT-5.5), and on AutomationBench-AA it debuted in first place. Newer Chinese models now score near the top on most broad benchmarks, including long-context tasks, multi-step coding, and tool coordination.

  2. Why it matters

    The Western lead has shrunk to three narrow areas—abstract pattern recognition (like ARC-AGI-2, where the gap widens to 89.2% vs. 60.4%), reliability in agentic tasks (Opus 5 passes 54% of repeated runs vs. K3's 39%), and offensive cybersecurity (where K3 scored 32% on ExploitBench vs. US leaders' ~76%). Below these frontiers, cheaper open models from China dominate. For investors betting on model exclusivity as a moat, this signals that raw performance alone no longer sustains a business—a concern Anthropic reportedly faces ahead of its IPO.

  3. What to watch

    The cybersecurity gap is shrinking fastest; GLM-5.3 (released August 14) scored 54.4% on ExploitBench, more than double its predecessor GLM-5.2, and even outpaced leading US models on CyberGym vulnerability detection. However, both Western labs and Chinese providers now restrict dangerous capabilities to controlled access (e.g., Anthropic's Mythos 5 through Project Glasswing, Z.ai limiting GLM-5.3's cyber functions to verified users with a two-week delay for safety work).

Ask the AI about this article →

Context & Analysis

The speed of Chinese lab advancement has shifted the competitive landscape fundamentally. A year and a half ago, when DeepSeek R1 shocked markets by competing with OpenAI's o1 at reportedly lower cost, the picture was murkier—Chinese models won on individual benchmarks (like AIME 2024) but trailed on others (like SimpleQA factual knowledge). By late June, the pattern began to change: Z.ai's GLM-5.2 still showed scattered strength, but the newest wave—K3, Qwen3.8-Max, GLM-5.3—revealed a genuine shift across nearly all broad, demanding evaluations.

The gap's collapse has created a strategic dilemma for Western labs. Measured by common benchmarks, what analysts once cited as a "few months" lead has evaporated in all but a handful of narrow domains. This matters because investors now face a hard question: if a freely downloadable model can replicate most paid capabilities within months, what business moat remains? Anthropic is reportedly fielding investor concerns ahead of its IPO by pointing to its remaining leadership at the very top tier. But below that peak, the field belongs largely to open, far cheaper Chinese models.

The distillation question looms over this shift. Both OpenAI and Anthropic allege that Chinese labs used their models at scale as teachers through API access—Anthropic documented over 16 million interactions across fraudulent accounts targeting agent reasoning, tool use, and reasoning traces. The timing of K3's release (shortly after Fable 5's launch) has prompted accusations that training windows were too tight, yet the same analyses show overlaps with the longer-available Opus 4.8. However, neither lab has published independently verifiable proof, and the technical constraints of black-box API access (weights, internal evaluations, and search processes remain hidden) mean distillation alone does not explain the breadth of convergence. Whether the advantage came from distillation or organic development, the conclusion is the same: a model lead only holds where the product isn't broadly offered, and no business can rest on that alone.

FAQ

Where do Chinese models still lag behind Western ones?
Three areas remain: on abstract pattern recognition tests like ARC-AGI-2 (Opus 5 at 89.2% vs. K3 at 60.4%), on reliability in agentic data analysis where pass^5 (solving 5 out of 5 independent runs) shows Opus 5 at 54% and K3 at 39%, and on offensive cyber capabilities where K3 scored 32% on ExploitBench versus US leaders averaging around 76%.
How did Chinese labs improve so quickly?
The body identifies distillation (using Western models as teachers) as a suspected mechanism. Anthropic reported that DeepSeek, Moonshot, and MiniMax ran over 16 million interactions through around 24,000 fraudulent accounts to access Western APIs; Moonshot's campaign alone involved 3.4 million interactions targeting agent reasoning and tool use. However, no lab has published verifiable evidence, and critics note the window between some Western model releases and K3's launch was too short to fully influence training.
Are Western labs restricting their strongest AI capabilities?
Yes. Anthropic's Mythos 5 (78% on ExploitBench) is available only under controlled conditions through Project Glasswing, while its public sibling Fable 5 remains at 40% because upstream safeguards blocked 407 of 410 test episodes. OpenAI uses similar restrictions with Daybreak and cyber variants. For the first time, Z.ai is adopting this pattern with GLM-5.3, delaying the weights release by about two weeks and limiting sensitive cyber functions to verified users.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleGrok leaks user data via encrypted malicious instructions

The AI news that matters, in one minute each morning.

Sign up free