AIToday
Large Language ModelsTHE DECODERPublished: Aug 7, 2026, 01:00 JST3 min read

Qwen3.8 Max matches Claude Opus 4.8 but trails Kimi K3 at higher cost

Qwen3.8 Max matches Claude Opus 4.8 but trails Kimi K3 at higher cost

Key takeaway

  • Alibaba's Qwen3.8 Max achieved performance parity with Claude Opus 4.8 on general benchmarks and surpassed Kimi K3 on work-related tasks, but the model relies on 64 reasoning steps instead of 14 and resends full conversation history repeatedly, raising its per-task cost to $1.14—more than double the prior version and higher than Kimi K3 at $0.86 for slightly better scores.

  • Hallucination rates also jumped significantly, from 23 to 40 percent.

3 Key Points

  1. What happened

    Alibaba's Qwen3.8 Max scored 56 on the Artificial Analysis Intelligence Index, a 10-point jump from Qwen3.7 Max (46), putting it on par with Claude Opus 4.8 but behind Kimi K3 (57). On the GDPval-AA benchmark for work-related tasks, Qwen3.8 Max reached 1,739 Elo points, surpassing Kimi K3 (1,685) but trailing Claude Opus 5 (1,852).

  2. Why it matters

    Despite lower token prices, Qwen3.8 Max costs more than double per task ($1.14 vs. $0.53 for the prior version) because it needs 64 steps instead of 14 and resends full conversation history at each step, making input tokens grow 15×. Kimi K3 delivers one point higher on the Intelligence Index for just $0.86 per task, offering better value. The model also shows regressions: hallucination rates jumped from 23 to 40 percent, and it guesses far more often instead of admitting uncertainty.

  3. What to watch

    Token pricing dropped (input from $2.50 to $2.00 per million tokens, output from $7.50 to $6.00 per million tokens, cache hits from $0.50 to $0.25), but the efficiency trade-off means cost-per-task remains a key competitive pressure against Kimi K3 and Claude Opus 5.

In Depth

Read the full story

Alibaba released Qwen3.8 Max on the Artificial Analysis Intelligence Index with a score of 56, marking a significant 10-point improvement over Qwen3.7 Max (46). According to Artificial Analysis, this puts Qwen3.8 Max on par with Claude Opus 4.8 and ahead of GLM-5.2 (51), but it falls short of Kimi K3 (57)—the higher-scoring alternative that also runs 25 percent cheaper.

On GDPval-AA, a benchmark designed to measure work-related task performance, Qwen3.8 Max jumped 468 Elo points to 1,739, surpassing Kimi K3 (1,685) but remaining below Claude Opus 5 (1,852). However, the method behind this improvement reveals a costly trade-off. The model achieves higher scores by employing 64 steps per task instead of 14, and it resends the full conversation history to the model at each step, causing input tokens to grow 15×.

This architectural choice has direct economic consequences. Although Alibaba reduced token prices—input dropped from $2.50 to $2.00 per million tokens, output from $7.50 to $6.00 per million tokens, and cache hits from $0.50 to $0.25—the volume increase means a single task on the Intelligence Index now costs $1.14, more than double Qwen3.7 Max ($0.53). By contrast, Kimi K3 scores one point higher at just $0.86 per task, and GLM-5.2 comes in at $0.57, making both substantially more cost-effective.

Qwen3.8 Max also exhibits concerning regressions. AA-LCR, which tests whether a model can correctly extract information from very long texts, dropped 2 points. AA-Omniscience, measuring accurate knowledge responses or honest admission of uncertainty, fell 10 points. The accuracy rate holds around 31 percent, but the hallucination rate jumped from 23 to 40 percent—the model guesses far more often instead of admitting it does not know, a behavioral shift that may undermine reliability despite surface-level performance gains.

FAQ

How does Qwen3.8 Max's performance compare to Kimi K3 and Claude Opus?
Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, matching Claude Opus 4.8 but trailing Kimi K3 (57). On GDPval-AA for work tasks, Qwen3.8 Max reaches 1,739 Elo, passing Kimi K3 (1,685) but below Claude Opus 5 (1,852).
Why is Qwen3.8 Max more expensive per task despite lower token prices?
Qwen3.8 Max requires 64 steps per task instead of 14, and resends the full conversation history at each step, causing input tokens to grow 15×. A single task now costs $1.14, more than double Qwen3.7 Max ($0.53), whereas Kimi K3 achieves a higher score for just $0.86 per task.
What performance regressions appear in Qwen3.8 Max?
AA-LCR dropped 2 points on long-text coherence, and AA-Omniscience fell 10 points on knowledge accuracy. Hallucination rate jumped from 23 to 40 percent, with the model guessing far more often instead of admitting it doesn't know.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleCTO Circle: Engineering Leaders Share AI-Native Organization Playbook

The AI news that matters, in one minute each morning.

Sign up free