[AINews] Qwen 3.8 Max(2.4T) and 27B, new open weights models for Coding and Cowork
Qwen 3.8 Max is a 2.4T-parameter flagship model from Alibaba, positioned as their most capable model to date, with open weights promised for release the following week alongside Qwen3.8-27B The model demonstrates exceptional long-horizon autonomous capabilities, including 10+ days of unattended coding, 125-hour autonomous AI research loops, and 500+ turns of chip design optimization In benchmark comparisons, it outperformed 87% of human teams (top 13% of 526 teams) in the WWW2025 Multimodal Dial
Analysis
TL;DR
- Qwen 3.8 Max is a 2.4T-parameter flagship model from Alibaba, positioned as their most capable model to date, with open weights promised for release the following week alongside Qwen3.8-27B
- The model demonstrates exceptional long-horizon autonomous capabilities, including 10+ days of unattended coding, 125-hour autonomous AI research loops, and 500+ turns of chip design optimization
- In benchmark comparisons, it outperformed 87% of human teams (top 13% of 526 teams) in the WWW2025 Multimodal Dialogue Intent Recognition Challenge within 24 hours, and achieved a 4.16x return in a 365-day e-commerce simulation
- API pricing is set at $2.00/M input tokens, $6.00/M output tokens, and $0.25/M cached tokens, with availability across Qwen Studio, API, Command Code, and Venice platforms
- This launch signals Alibaba's renewed commitment to the open-weight ecosystem following the "Qwen Exodus" and management changes, directly countering skepticism about their open model strategy
Why It Matters
Qwen 3.8 Max represents a significant milestone for the open-weight model movement, demonstrating that a 2.4T-parameter model can compete at the highest levels while remaining accessible to the broader research community. For AI practitioners, the model's native multimodal execution loop—where vision is integrated into the agent's reasoning pipeline rather than serving as a passive input—signals a shift toward more capable autonomous agents capable of sustained, real-world task execution across coding, research, hardware design, and business operations.
Technical Details
- Model Architecture: 2.4T-parameter flagship model (Qwen3.8-Max) with a companion 27B-parameter open-weight variant (Qwen3.8-27B), both featuring native multimodal intelligence where visual feedback is embedded directly into planning, coding, and GUI interaction loops
- Autonomous Coding & Research: Demonstrated 10+ days of unattended self-evolving coding harness development, 125-hour autonomous research loops that rebuilt and improved upon the "Unified Data Selection for LLM Reasoning" pipeline (achieving +2.71 points over the original benchmark), and complete silicon design flows reducing gate count from 8,298 to 678 with 81% die area reduction at 500 MHz timing closure
- Multimodal Agent Framework: Released Qwen-MM-Plugins to extend multimodal capabilities into existing agent frameworks, enabling cross-platform application recreation (desktop, mobile, web) with native visual feedback across the execution loop
- E-Commerce Performance: Achieved a 4.16x return (¥416,252 balance) in the E-Commerce Bench—a 365-day store operation simulation—through continuous game-theoretic negotiation and inventory planning, outperforming competing models
- Production Workflows: Validated across hundreds of professional workflows including corporate legal reviews, UI/UX design, structural engineering models, and automated ETF quant research, with infrastructure support from Baseten, Hermes Agent, and Command Code
Industry Insight
- The open-weight commitment for a 2.4T model challenges the prevailing trend toward increasingly closed API-only offerings, potentially accelerating the open-source model ecosystem and giving smaller teams access to frontier-capable architectures
- The integration of vision as an active execution component rather than a passive input channel represents a meaningful architectural evolution for agentic systems, suggesting that future benchmark evaluations will increasingly prioritize sustained multi-step autonomous performance over static accuracy metrics
- Alibaba's strategic pivot back toward open weights following internal turbulence signals that even companies facing management changes may recognize open-source as a competitive differentiator, potentially encouraging other labs to maintain or expand their open-weight commitments despite commercial pressures
Disclaimer: The above content is generated by AI and is for reference only.