AIToday
Open-Source AITHE DECODERPublished: Aug 3, 2026, 22:01 JST2 min read

Alibaba releases Qwen3.8-Max, 2.4T-parameter AI for week-long autonomous tasks

Alibaba releases Qwen3.8-Max, 2.4T-parameter AI for week-long autonomous tasks

Key takeaway

  • Alibaba released Qwen3.8-Max, a 2.4 trillion parameter AI model designed to work autonomously on complex tasks over days or weeks without human intervention.

  • The model demonstrated the ability to reproduce research papers, design chips more efficiently than prior attempts, and win competitions against human teams, while also mastering multimodal tasks.

  • Weights will be publicly available next week, making it the first open-weight model in Alibaba's flagship Qwen-Max class.

3 Key Points

  1. What happened

    Alibaba's Qwen team unveiled Qwen3.8-Max, a 2.4 trillion parameter language model with 95 billion active parameters per query, designed to complete complex tasks independently over extended periods. The weights will be made publicly available next week on Hugging Face and ModelScope.

  2. Why it matters

    The model demonstrates sustained autonomous capability across diverse domains—it reproduced research results and beat published benchmarks by 2.7 points on AIME24 math, solved a chip-design problem by reducing gate count from 8,298 to 678, and outperformed 458 of 526 human teams in a competition. For developers and researchers, this is the first open-weight model in the Qwen-Max class, removing a barrier to experimentation with long-horizon AI agents.

  3. What to watch

    Qwen3.8-Max is available now through QwenCloud; weights go live next week on Hugging Face and ModelScope. The model supports OpenAI Chat Completions and Anthropic API formats, and includes a reasoning_effort parameter with three speed-versus-thoroughness levels. Independent verification of the benchmarks is still pending.

Ask the AI about this article →

Context & Analysis

Qwen3.8-Max represents a significant milestone in open-weight AI development, moving beyond single-turn interactions to sustained, multi-day autonomous work. The model's architecture builds on Qwen3.5 but scales to 2.4 trillion total parameters with substantially expanded reinforcement learning training covering multi-day workflows, nested structures, and multiple agent harnesses—a departure from the typical single-task focus. This shift in training methodology appears to be the foundation for its long-horizon capability: the team's internal benchmark index rose from 0.474 to 0.725, with performance peaking around 4,000 training environments before plateauing slightly.

The concrete case studies reveal the model's range. In competitive terms, it outperformed 458 of 526 human teams in a multimodal dialogue challenge within 24 hours and beat published research by 2.7 points on the AIME24 math benchmark. In engineering contexts, it achieved an 81 percent reduction in physical chip area through iterative gate-count optimization. The e-commerce simulation—where the model quadrupled starting capital and outperformed its predecessor by more than 2.5× and a rival (GLM 5.2) by 38 percent—demonstrates planning and negotiation across hundreds of interaction rounds. These results are self-reported and not yet independently verified, though the breadth of demonstrated capabilities suggests the model moves beyond typical benchmark inflation.

FAQ

When will the weights be available to the public?
The weights are set to go live next week on Hugging Face and ModelScope, according to Alibaba.
What are some examples of long-horizon tasks the model can perform?
Qwen3.8-Max spent 16 days building the command-line tool oh-my-cli with 265 commits and 127 pull requests, reproduced a research paper on data selection over five days by writing 7,600 lines of code and running 33 GPU training jobs, reduced a cryptographic chip design from 8,298 gates to 678 over roughly 500 iterations, and ran a full fiscal year e-commerce simulation starting with 100,000 yuan and ending with 416,252 yuan.
What APIs does Qwen3.8-Max support?
The model supports both OpenAI's Chat Completions format and Anthropic's API protocol, allowing it to work directly with Claude Code, Codex, Qoder CLI, Qwen Code, and OpenClaw.

Get the latest Open-Source AI news every morning

For example, today's edition would include:

  • DataAgent launches with $10M to auto-fix Kubernetes faultsSiliconANGLE AI · 1h ago
  • Z.ai runs GLM on 100,000 Chinese AI chipsDIGITIMES Asia · 4h ago
  • Broadcom Unveils VMware AI Factory for Faster Private AITop Companies AI · 14h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAltman pitches AI podcast for kids' schedules; backlash follows