
Alibaba's Qwen announced Qwen3.8-Max, a 2.4T-parameter model designed for coding and autonomous work, with open weights promised next week at $2 input / $6 output per million tokens on API.
Early benchmarks rank it #2 to #4 on major coding and vision leaderboards, with 87.3% on SWE-bench, matching Claude Opus 4.7's cost-adjusted score.
However, deployment requires at least 8 high-end GPUs, and users have raised concerns about potential geographic license restrictions covering the US, EU, UK, and Korea.
What happened
Alibaba announced Qwen3.8-Max, a 2.4T-parameter model focused on coding and autonomous agent work, and committed to releasing open weights next week alongside a smaller Qwen3.8-27B model. The API is priced at $2 input / $6 output per million tokens. Early third-party benchmarks placed it #4 in Frontend Code Arena (1,668 Elo), #2 in Vision Arena (1,305 Elo), and #2 among open-weight models on Vals Index (66.1 score).
Why it matters
The release signals that Chinese open-weight models are now competing directly with top Western closed models in coding, vision, and agent-driven tasks. The API price is lower than Qwen3.7-Max ($2.50/$7.50), and independent testing showed the model matched Claude Opus 4.7 on cost-normalized benchmarks while achieving 87.3% on SWE-bench. However, the 2.4T sparse architecture is not practical for local inference—it requires 8+ high-end GPUs—so real-world adoption may depend more on the smaller 27B descendant.
What to watch
Open weights release is promised for next week. A key caveat: users flagged potential geographic restrictions in the license terms that may prohibit download and use in the US, EU, UK, and Korea; Alibaba has not publicly clarified the licensing terms. Practical ecosystem impact may hinge on the 27B model's availability and capabilities, not just the 2.4T flagship's benchmark rank.
Alibaba Qwen announced Qwen3.8-Max on August 3, 2026, positioning it as the company's "most capable model to date" and a flagship focused on autonomous coding, long-horizon agent work, and native multimodal reasoning. The model is a 2.4T-parameter sparse architecture with approximately 95B active parameters per token, yielding a ~4% mixture-of-experts activation ratio. It supports a 1M-token context window and up to 128k tokens of output.
The API pricing is $2.00 per million input tokens, $6.00 per million output tokens, and $0.25 per million cached tokens—a reduction from Qwen3.7-Max pricing ($2.50 / $7.50). Alibaba committed to releasing open weights for both Qwen3.8-Max and Qwen3.8-27B "next week," making the flagship model's weights freely available rather than API-only.
Alibaba's own claims for the model are notably aggressive: 10+ days of autonomous coding with a public GitHub trace, 500+ turns for chip design optimization, 365 days of e-commerce strategy execution, and native multimodal planning where vision is integrated into the execution loop rather than serving merely as an input channel. The company simultaneously pushed availability across multiple surfaces—Qwen Studio, API, Command Code, and Venice—with rapid partner support from Baseten, Hermes Agent, and Command Code.
Third-party benchmarks confirmed strong placements. Arena reported Qwen3.8-Max at #4 in Frontend Code Arena with 1,668 Elo (trailing Claude Opus 5 Max at 1,705 and Kimi K3 Max at 1,676), and #2 in Vision Arena at 1,305 Elo, 13 points behind Claude Fable 5 High. Vals AI ranked it #2 among open-weight models and #10 overall out of 43 models with a score of 66.1, matching Claude Opus 4.7 (also 66.1) while costing roughly 2.3× less per test ($2.68 vs. $6.17). On SWE-bench, Qwen3.8-Max achieved 87.3%, ahead of GPT-5.5 (82.6%) and GLM-5.2 (83.3%), though behind Claude Opus 4.8 (89.2%). Terminal-Bench 2.1 showed 67.4, up from 61.0 for Qwen3.7-Max, representing an 8.6-point gain over ~2.5 months.
However, important caveats emerged in the reaction. Jamin Ball emphasized that pricing comparisons are misleading because vanilla token prices ignore token efficiency and because these models are operationally prohibitive for local deployment: the 2.4T Qwen requires at least 8 H100/B200 GPUs and over 1TB of memory just to load weights, effectively identical to Kimi K3's infrastructure demands. Users flagged a more serious concern: potential geographic restrictions in the license terms that appear to cover the US, EU, UK, and Korea, with one observer (OstrisAI) reading the terms as forbidding even download from the US. Alibaba did not issue a clarifying tweet on licensing within the dataset, leaving the "open weights" claim ambiguous—the model may be restricted from commercial deployment, export, or even download in major regions.
The release was widely interpreted as a strategic pivot by Alibaba away from API-only exclusivity toward ecosystem influence, driven by competition from DeepSeek, Kimi, and other Chinese open models that had already weakened the premium of keeping top-tier systems closed. The real adoption wave, observers suggested, may come not from the 2.4T flagship but from the smaller Qwen3.8-27B, which practitioners expect to inherit post-training or distilled capabilities while remaining practical for developers on consumer hardware.
Alibaba's Qwen3.8-Max release represents a strategic shift in how the company competes in the open-model space. After the Qwen leadership exodus last year, there was genuine doubt about whether the lab would continue releasing cutting-edge open models or focus solely on proprietary APIs. This announcement signals a deliberate ecosystem play: by open-sourcing a flagship 2.4T model alongside a smaller 27B sibling, Alibaba is competing not just on model quality but on community adoption and downstream ecosystem influence—a pattern driven by competing open-model releases from DeepSeek, Kimi K3, and others that weakened the premium of keeping top-tier systems API-only.
The model's benchmark performance is genuinely strong. It debuted at #4 in Frontend Code Arena (1,668 Elo, trailing only Claude Opus 5 Max and Kimi K3), #2 in Vision Arena, and achieved 87.3% on SWE-bench, outperforming GPT-5.5 (82.6%) and matching Claude Opus 4.7 on cost-normalized testing despite being offered at lower raw token pricing. The autonomous coding demonstrations—10+ days unattended, 500+ turns of chip design, 365 days of e-commerce simulation—reflect genuine post-training and agentic capability, though these are vendor-reported claims rather than independently validated benchmarks.
However, two practical constraints temper the immediate impact. First, the 2.4T sparse model is not locally deployable—it requires 8+ H100/B200 GPUs and over 1TB of memory, placing it in the same operational class as Kimi K3 and GLM-5.2. Real adoption may flow instead from the smaller Qwen3.8-27B, which inherits distilled capabilities but runs on consumer hardware. Second, users flagged potential geographic license restrictions covering the US, EU, UK, and Korea; without clarification from Alibaba, the "open weights" label carries ambiguity about whether the model can legally be downloaded or commercialized in key markets. This licensing uncertainty, echoed in parallel discussions around MiniMax H3, represents a meaningful gap between the marketing claim and engineering reality.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Silicon Motion announced a private placement of US$1 billion in aggregate principal amount of 0% convertible s…

Sundar Pichai, CEO of Alphabet and Google, announced on August 11 that the Gemini app's monthly active users (…

A new platform called frontier.fast has launched an open competition where anyone can submit code patches to m…

Researchers published findings showing AMIE (Video), a Gemini-based AI system, performed at or above the level…

An AI system generated a research draft that strengthened a mathematical bound related to the Riemann hypothes…

After former lead writer Stella Sacco posted on Bluesky that Saber replaced her with ChatGPT midway through de…

The AI news that matters, in one minute each morning.
Sign up free