Alibaba's open-weight Qwen3.8-Max takes on long-horizon AI tasks with 2.4 trillion parameters
Alibaba unveiled Qwen3.8-Max, a 2.4-trillion-parameter open-weight language model with 95 billion active parameters per query, built on the Qwen3.5 architecture The model demonstrates exceptional long-horizon autonomous task completion, including a 16-day software build, 5-day research reproduction, and 24-hour competition entry without human intervention Case studies show the model reducing chip design from 8,298 to 678 logic gates (81% area reduction) and quadrupling simulated capital in an e-
Analysis
TL;DR
- Alibaba unveiled Qwen3.8-Max, a 2.4-trillion-parameter open-weight language model with 95 billion active parameters per query, built on the Qwen3.5 architecture
- The model demonstrates exceptional long-horizon autonomous task completion, including a 16-day software build, 5-day research reproduction, and 24-hour competition entry without human intervention
- Case studies show the model reducing chip design from 8,298 to 678 logic gates (81% area reduction) and quadrupling simulated capital in an e-commerce benchmark, outperforming GLM 5.2 by 38%
- Qwen3.8-Max is the first Qwen-Max class model with publicly released weights, competing on par with Claude Opus 4.8, Fable 5, and GPT-5.6 Sol on internal benchmarks
- New capabilities include processing 200+ page documents and 100+ hour videos, plus the introduction of RecreationBench for app reconstruction without source code and Qwen-MM-Plugins for multimodal agent extensions
Why It Matters
Qwen3.8-Max represents a significant milestone in autonomous AI agents capable of sustaining complex, multi-day workflows without human oversight—a critical step toward practical agentic AI systems. The open-weight release of a model in this capability tier challenges the closed-model dominance and could accelerate research into long-horizon reasoning, autonomous coding, and multi-step planning. For practitioners, it demonstrates that open-weight models can now compete with frontier proprietary systems on agentic and reasoning benchmarks.
Technical Details
- Architecture: 2.4T total parameters, 95B active per query, built on Qwen3.5 architecture with MoE (Mixture of Experts) design
- Autonomous coding: Completed 16-day oh-my-cli project with 265 commits, 127 PRs, and 151 issues; reproduced and improved research paper results over 5 days using 125 GPU-hours across 33 training jobs
- Chip design optimization: Reduced cryptographic circuit from 8,298 to 678 gates through ~500 iterations, achieving 81% physical area reduction (106x106 to 46x46 micrometers) after OpenROAD layout
- Multimodal & benchmarks: Handles 200+ page documents and 100+ hour videos; PaperBench score of 93 (highest reported); TerminalBench 2.1 score of 86.6; RecreationBench tests app reconstruction via interaction-only observation across Ubuntu, macOS, Windows, Android, and web
- E-commerce simulation: Quadrupled 100,000 yuan capital to 416,252 yuan in a full fiscal year simulation with 152 hidden scammers, outperforming GLM 5.2 by 38%
Industry Insight
- The open-weight release of a frontier-capable model signals intensifying competition between Chinese and Western AI labs, potentially compressing the performance gap in agentic AI workflows
- Long-horizon autonomous task completion is emerging as the key differentiator for next-generation AI systems; practitioners should evaluate these models for complex, multi-step workflows rather than single-turn tasks
- The introduction of benchmarks like RecreationBench and E-Commerce-Bench reflects a broader industry shift toward evaluating AI on sustained, real-world agentic performance rather than static benchmark scores
Disclaimer: The above content is generated by AI and is for reference only.