AIToday
Large Language ModelsTHE DECODERPublished: Aug 7, 2026, 01:00 JST

Qwen3.8 Max matches Claude Opus 4.8 but trails Kimi K3 at higher cost

Qwen3.8 Max matches Claude Opus 4.8 but trails Kimi K3 at higher cost

3 Key Points

  1. What happened

    Alibaba's Qwen3.8 Max scored 56 on the Artificial Analysis Intelligence Index, a 10-point jump from Qwen3.7 Max (46), putting it on par with Claude Opus 4.8 but behind Kimi K3 (57). On the GDPval-AA benchmark for work-related tasks, Qwen3.8 Max reached 1,739 Elo points, surpassing Kimi K3 (1,685) but trailing Claude Opus 5 (1,852).

  2. Why it matters

    Despite lower token prices, Qwen3.8 Max costs more than double per task ($1.14 vs. $0.53 for the prior version) because it needs 64 steps instead of 14 and resends full conversation history at each step, making input tokens grow 15×. Kimi K3 delivers one point higher on the Intelligence Index for just $0.86 per task, offering better value. The model also shows regressions: hallucination rates jumped from 23 to 40 percent, and it guesses far more often instead of admitting uncertainty.

  3. What to watch

    Token pricing dropped (input from $2.50 to $2.00 per million tokens, output from $7.50 to $6.00 per million tokens, cache hits from $0.50 to $0.25), but the efficiency trade-off means cost-per-task remains a key competitive pressure against Kimi K3 and Claude Opus 5.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

FAQ
How does Qwen3.8 Max's performance compare to Kimi K3 and Claude Opus?
Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, matching Claude Opus 4.8 but trailing Kimi K3 (57). On GDPval-AA for work tasks, Qwen3.8 Max reaches 1,739 Elo, passing Kimi K3 (1,685) but below Claude Opus 5 (1,852).
Why is Qwen3.8 Max more expensive per task despite lower token prices?
Qwen3.8 Max requires 64 steps per task instead of 14, and resends the full conversation history at each step, causing input tokens to grow 15×. A single task now costs $1.14, more than double Qwen3.7 Max ($0.53), whereas Kimi K3 achieves a higher score for just $0.86 per task.
What performance regressions appear in Qwen3.8 Max?
AA-LCR dropped 2 points on long-text coherence, and AA-Omniscience fell 10 points on knowledge accuracy. Hallucination rate jumped from 23 to 40 percent, with the model guessing far more often instead of admitting it doesn't know.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Anthropic ships Claude Sonnet 5.5, 30%+ faster at same priceITmedia AI+ · 1h ago
  • Dentsu AI clears 99% on 20万件超 approvalsITmedia AI+ · 1h ago
  • Nvidia launches tool to quarantine rogue AI agentsSemafor Tech · 1h ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleCTO Circle: Engineering Leaders Share AI-Native Organization Playbook