
What happened
Alibaba's Qwen3.8 Max scored 56 on the Artificial Analysis Intelligence Index, a 10-point jump from Qwen3.7 Max (46), putting it on par with Claude Opus 4.8 but behind Kimi K3 (57). On the GDPval-AA benchmark for work-related tasks, Qwen3.8 Max reached 1,739 Elo points, surpassing Kimi K3 (1,685) but trailing Claude Opus 5 (1,852).
Why it matters
Despite lower token prices, Qwen3.8 Max costs more than double per task ($1.14 vs. $0.53 for the prior version) because it needs 64 steps instead of 14 and resends full conversation history at each step, making input tokens grow 15×. Kimi K3 delivers one point higher on the Intelligence Index for just $0.86 per task, offering better value. The model also shows regressions: hallucination rates jumped from 23 to 40 percent, and it guesses far more often instead of admitting uncertainty.
What to watch
Token pricing dropped (input from $2.50 to $2.00 per million tokens, output from $7.50 to $6.00 per million tokens, cache hits from $0.50 to $0.25), but the efficiency trade-off means cost-per-task remains a key competitive pressure against Kimi K3 and Claude Opus 5.
Summaries like this, in your inbox every morning.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Anthropic released Claude Sonnet 5.5 on September 28, the second model in its Claude 5.5 family, calling it fa…

Dentsu Group embedded generative AI in its accounting approval process and says it reached 99% accuracy on the…

Nvidia launched a platform Monday that it says can monitor AI agents and quarantine suspicious ones, following…

A September 2026 guide organizes AI coding tools into three types — terminal/agent tools like Claude Code, AI-…

On September 25, MAGI Games' three debating AIs split 1-1-1 on three mini-game proposals and refused to budge…

Running Claude Code in one shared working tree produced 8 accidents and near-misses between June and September…
