AI News AI资讯 5h ago Updated 2h ago 更新于 2小时前 50

[AINews] Qwen 3.8 Max(2.4T) and 27B, new open weights models for Coding and Cowork 【AI新闻】Qwen 3.8 Max(2.4T)和27B,面向编程与协作的全新开源权重模型

Qwen 3.8 Max is a 2.4T-parameter flagship model from Alibaba, positioned as their most capable model to date, with open weights promised for release the following week alongside Qwen3.8-27B The model demonstrates exceptional long-horizon autonomous capabilities, including 10+ days of unattended coding, 125-hour autonomous AI research loops, and 500+ turns of chip design optimization In benchmark comparisons, it outperformed 87% of human teams (top 13% of 526 teams) in the WWW2025 Multimodal Dial Qwen 3.8 Max是阿里巴巴的2.4T参数旗舰模型,定位为迄今为止最强大的模型,承诺在下周与Qwen3.8-27B一同发布开源权重 该模型展现出卓越的长周期自主能力,包括超过10天的无人值守编程、125小时的自主AI研究循环,以及500多轮的芯片设计优化 在基准测试中,它在24小时内于WWW2025多模态对话意图识别挑战赛中超越了87%的人类团队(526支团队中的前13%),并在365天的电商模拟中实现了4.16倍的回报 API定价为输入token每百万2.00美元、输出token每百万6.00美元、缓存token每百万0.25美元,将在Qwen Studio、API、Command C

72
Hot 热度
68
Quality 质量
75
Impact 影响力

Analysis 深度分析

TL;DR

  • Qwen 3.8 Max is a 2.4T-parameter flagship model from Alibaba, positioned as their most capable model to date, with open weights promised for release the following week alongside Qwen3.8-27B
  • The model demonstrates exceptional long-horizon autonomous capabilities, including 10+ days of unattended coding, 125-hour autonomous AI research loops, and 500+ turns of chip design optimization
  • In benchmark comparisons, it outperformed 87% of human teams (top 13% of 526 teams) in the WWW2025 Multimodal Dialogue Intent Recognition Challenge within 24 hours, and achieved a 4.16x return in a 365-day e-commerce simulation
  • API pricing is set at $2.00/M input tokens, $6.00/M output tokens, and $0.25/M cached tokens, with availability across Qwen Studio, API, Command Code, and Venice platforms
  • This launch signals Alibaba's renewed commitment to the open-weight ecosystem following the "Qwen Exodus" and management changes, directly countering skepticism about their open model strategy

Why It Matters

Qwen 3.8 Max represents a significant milestone for the open-weight model movement, demonstrating that a 2.4T-parameter model can compete at the highest levels while remaining accessible to the broader research community. For AI practitioners, the model's native multimodal execution loop—where vision is integrated into the agent's reasoning pipeline rather than serving as a passive input—signals a shift toward more capable autonomous agents capable of sustained, real-world task execution across coding, research, hardware design, and business operations.

Technical Details

  • Model Architecture: 2.4T-parameter flagship model (Qwen3.8-Max) with a companion 27B-parameter open-weight variant (Qwen3.8-27B), both featuring native multimodal intelligence where visual feedback is embedded directly into planning, coding, and GUI interaction loops
  • Autonomous Coding & Research: Demonstrated 10+ days of unattended self-evolving coding harness development, 125-hour autonomous research loops that rebuilt and improved upon the "Unified Data Selection for LLM Reasoning" pipeline (achieving +2.71 points over the original benchmark), and complete silicon design flows reducing gate count from 8,298 to 678 with 81% die area reduction at 500 MHz timing closure
  • Multimodal Agent Framework: Released Qwen-MM-Plugins to extend multimodal capabilities into existing agent frameworks, enabling cross-platform application recreation (desktop, mobile, web) with native visual feedback across the execution loop
  • E-Commerce Performance: Achieved a 4.16x return (¥416,252 balance) in the E-Commerce Bench—a 365-day store operation simulation—through continuous game-theoretic negotiation and inventory planning, outperforming competing models
  • Production Workflows: Validated across hundreds of professional workflows including corporate legal reviews, UI/UX design, structural engineering models, and automated ETF quant research, with infrastructure support from Baseten, Hermes Agent, and Command Code

Industry Insight

  • The open-weight commitment for a 2.4T model challenges the prevailing trend toward increasingly closed API-only offerings, potentially accelerating the open-source model ecosystem and giving smaller teams access to frontier-capable architectures
  • The integration of vision as an active execution component rather than a passive input channel represents a meaningful architectural evolution for agentic systems, suggesting that future benchmark evaluations will increasingly prioritize sustained multi-step autonomous performance over static accuracy metrics
  • Alibaba's strategic pivot back toward open weights following internal turbulence signals that even companies facing management changes may recognize open-source as a competitive differentiator, potentially encouraging other labs to maintain or expand their open-weight commitments despite commercial pressures

摘要

Qwen 3.8 Max是阿里巴巴的2.4T参数旗舰模型,定位为迄今为止最强大的模型,承诺在下周与Qwen3.8-27B一同发布开源权重
该模型展现出卓越的长周期自主能力,包括超过10天的无人值守编程、125小时的自主AI研究循环,以及500多轮的芯片设计优化
在基准测试中,它在24小时内于WWW2025多模态对话意图识别挑战赛中超越了87%的人类团队(526支团队中的前13%),并在365天的电商模拟中实现了4.16倍的回报
API定价为输入token每百万2.00美元、输出token每百万6.00美元、缓存token每百万0.25美元,将在Qwen Studio、API、Command Code和Venice平台上线
此次发布标志着阿里巴巴在"Qwen大撤离"和管理层变动后重新承诺开源权重生态,直接回应了外界对其开源模型战略的质疑

深度分析

一句话总结

  • Qwen 3.8 Max是阿里巴巴的2.4T参数旗舰模型,定位为迄今为止最强大的模型,承诺在下周与Qwen3.8-27B一同发布开源权重
  • 该模型展现出卓越的长周期自主能力,包括超过10天的无人值守编程、125小时的自主AI研究循环,以及500多轮的芯片设计优化
  • 在基准测试中,它在24小时内于WWW2025多模态对话意图识别挑战赛中超越了87%的人类团队(526支团队中的前13%),并在365天的电商模拟中实现了4.16倍的回报
  • API定价为输入token每百万2.00美元、输出token每百万6.00美元、缓存token每百万0.25美元,将在Qwen Studio、API、Command Code和Venice平台上线
  • 此次发布标志着阿里巴巴在"Qwen大撤离"和管理层变动后重新承诺开源权重生态,直接回应了外界对其开源模型战略的质疑

为何重要

Qwen 3.8 Max代表了开源权重模型运动的一个重要里程碑,证明了一个2.4T参数的模型可以在最高水平上竞争,同时保持对更广泛研究社区的可用性。对于AI从业者来说,该模型的原生多模态执行循环——视觉被集成到代理的推理管道中,而非作为被动输入——标志着向更强大的自主代理的转变,这些代理能够持续地进行真实世界的工作

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Open Source 开源 LLM 大模型 Code Generation 代码生成 Product Launch 产品发布 Training 训练