
Alibaba released Qwen3.8-Max, a 2.4 trillion parameter AI model designed to work autonomously on complex tasks over days or weeks without human intervention.
The model demonstrated the ability to reproduce research papers, design chips more efficiently than prior attempts, and win competitions against human teams, while also mastering multimodal tasks.
Weights will be publicly available next week, making it the first open-weight model in Alibaba's flagship Qwen-Max class.
What happened
Alibaba's Qwen team unveiled Qwen3.8-Max, a 2.4 trillion parameter language model with 95 billion active parameters per query, designed to complete complex tasks independently over extended periods. The weights will be made publicly available next week on Hugging Face and ModelScope.
Why it matters
The model demonstrates sustained autonomous capability across diverse domains—it reproduced research results and beat published benchmarks by 2.7 points on AIME24 math, solved a chip-design problem by reducing gate count from 8,298 to 678, and outperformed 458 of 526 human teams in a competition. For developers and researchers, this is the first open-weight model in the Qwen-Max class, removing a barrier to experimentation with long-horizon AI agents.
What to watch
Qwen3.8-Max is available now through QwenCloud; weights go live next week on Hugging Face and ModelScope. The model supports OpenAI Chat Completions and Anthropic API formats, and includes a reasoning_effort parameter with three speed-versus-thoroughness levels. Independent verification of the benchmarks is still pending.
Ask the AI about this article →
Qwen3.8-Max represents a significant milestone in open-weight AI development, moving beyond single-turn interactions to sustained, multi-day autonomous work. The model's architecture builds on Qwen3.5 but scales to 2.4 trillion total parameters with substantially expanded reinforcement learning training covering multi-day workflows, nested structures, and multiple agent harnesses—a departure from the typical single-task focus. This shift in training methodology appears to be the foundation for its long-horizon capability: the team's internal benchmark index rose from 0.474 to 0.725, with performance peaking around 4,000 training environments before plateauing slightly.
The concrete case studies reveal the model's range. In competitive terms, it outperformed 458 of 526 human teams in a multimodal dialogue challenge within 24 hours and beat published research by 2.7 points on the AIME24 math benchmark. In engineering contexts, it achieved an 81 percent reduction in physical chip area through iterative gate-count optimization. The e-commerce simulation—where the model quadrupled starting capital and outperformed its predecessor by more than 2.5× and a rival (GLM 5.2) by 38 percent—demonstrates planning and negotiation across hundreds of interaction rounds. These results are self-reported and not yet independently verified, though the breadth of demonstrated capabilities suggests the model moves beyond typical benchmark inflation.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Israeli startup DataAgent Ltd
Chinese large-model developer Z.ai says it can now support large-scale inference using roughly 100,000 domesti…

Broadcom announced VMware AI Factory, a software-defined foundation for VMware Private AI Cloud, at VMware Exp…

OpenClaw launched version 2.0, its largest update yet, with a version number of 2026.8.1
David Heinemeier Hansson (DHH), creator of Ruby on Rails, has released Omarchy 4.0 (Omarchy Quattro), the late…

Debian voted to allow developers to use AI tools in contributions to the Linux distribution, covering developm…
