
Alibaba's Qwen3.8 Max achieved performance parity with Claude Opus 4.8 on general benchmarks and surpassed Kimi K3 on work-related tasks, but the model relies on 64 reasoning steps instead of 14 and resends full conversation history repeatedly, raising its per-task cost to $1.14—more than double the prior version and higher than Kimi K3 at $0.86 for slightly better scores.
Hallucination rates also jumped significantly, from 23 to 40 percent.
What happened
Alibaba's Qwen3.8 Max scored 56 on the Artificial Analysis Intelligence Index, a 10-point jump from Qwen3.7 Max (46), putting it on par with Claude Opus 4.8 but behind Kimi K3 (57). On the GDPval-AA benchmark for work-related tasks, Qwen3.8 Max reached 1,739 Elo points, surpassing Kimi K3 (1,685) but trailing Claude Opus 5 (1,852).
Why it matters
Despite lower token prices, Qwen3.8 Max costs more than double per task ($1.14 vs. $0.53 for the prior version) because it needs 64 steps instead of 14 and resends full conversation history at each step, making input tokens grow 15×. Kimi K3 delivers one point higher on the Intelligence Index for just $0.86 per task, offering better value. The model also shows regressions: hallucination rates jumped from 23 to 40 percent, and it guesses far more often instead of admitting uncertainty.
What to watch
Token pricing dropped (input from $2.50 to $2.00 per million tokens, output from $7.50 to $6.00 per million tokens, cache hits from $0.50 to $0.25), but the efficiency trade-off means cost-per-task remains a key competitive pressure against Kimi K3 and Claude Opus 5.
Alibaba released Qwen3.8 Max on the Artificial Analysis Intelligence Index with a score of 56, marking a significant 10-point improvement over Qwen3.7 Max (46). According to Artificial Analysis, this puts Qwen3.8 Max on par with Claude Opus 4.8 and ahead of GLM-5.2 (51), but it falls short of Kimi K3 (57)—the higher-scoring alternative that also runs 25 percent cheaper.
On GDPval-AA, a benchmark designed to measure work-related task performance, Qwen3.8 Max jumped 468 Elo points to 1,739, surpassing Kimi K3 (1,685) but remaining below Claude Opus 5 (1,852). However, the method behind this improvement reveals a costly trade-off. The model achieves higher scores by employing 64 steps per task instead of 14, and it resends the full conversation history to the model at each step, causing input tokens to grow 15×.
This architectural choice has direct economic consequences. Although Alibaba reduced token prices—input dropped from $2.50 to $2.00 per million tokens, output from $7.50 to $6.00 per million tokens, and cache hits from $0.50 to $0.25—the volume increase means a single task on the Intelligence Index now costs $1.14, more than double Qwen3.7 Max ($0.53). By contrast, Kimi K3 scores one point higher at just $0.86 per task, and GLM-5.2 comes in at $0.57, making both substantially more cost-effective.
Qwen3.8 Max also exhibits concerning regressions. AA-LCR, which tests whether a model can correctly extract information from very long texts, dropped 2 points. AA-Omniscience, measuring accurate knowledge responses or honest admission of uncertainty, fell 10 points. The accuracy rate holds around 31 percent, but the hallucination rate jumped from 23 to 40 percent—the model guesses far more often instead of admitting it does not know, a behavioral shift that may undermine reliability despite surface-level performance gains.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Security researchers led by Alexander Panfilov discovered a vulnerability in the APIs of all major AI provider…

CEO Sundar Pichai announced that the Gemini app has surpassed 1 billion monthly active users, making it the 14…

River AI, founded by xAI co-founder Igor Babuschkin, raised $1.1 billion in a seed/Series A round led by Gener…

An unreleased Anthropic model significantly increased the lower bound of solutions for which the Riemann hypot…

Google Research and Google DeepMind have advanced AMIE, a research medical AI system built on Gemini and Proje…

ONESTRUCTION, Inc. built Ishigaki-IDS, a foundation model specialized for construction industry BIM workflows…

The AI news that matters, in one minute each morning.
Sign up free