AI News AI资讯 1d ago Updated 22h ago 更新于 22小时前 69

Latest open artifacts (#23): Laguna S2.1, Inkling, & Kimi K3 show the utility of open models on the Pareto frontier 最新开源模型(第23期):Laguna S2.1、Inkling 和 Kimi K3 展现 Pareto 前沿上开源模型的实用性

Industry consolidation predicted for 2026-2027 has not materialized; instead, more companies are investing hundreds of millions to billions in training strong open models Token demand is surging as models become more efficient, making "building token machines" a recognized path to value for labs Thinking Machines emerged as an unexpected open-model leader, with their finetuning service generating hundreds of millions in annual revenue while releasing top U.S. open-weight models Chinese labs main 预测的AI实验室整合并未如期发生,更多公司仍在投入数亿至数十亿美元训练模型并开源发布 Token需求持续高涨,"构建token机器"成为AI价值新路径,开源微调服务可带来数亿美元年收入 中国实验室(小米、美团等)保持强劲节奏,开源模型市场份额争夺进入关键时期 许可证策略成为战略变量:Apache 2.0、OpenMDW与非商业许可证各有利弊,影响商业与政策格局 多模态MoE架构成为主流,参数规模从百亿到万亿级并行发展,性价比与可部署性并重

65
Hot 热度
70
Quality 质量
68
Impact 影响力

Analysis 深度分析

TL;DR

  • Industry consolidation predicted for 2026-2027 has not materialized; instead, more companies are investing hundreds of millions to billions in training strong open models
  • Token demand is surging as models become more efficient, making "building token machines" a recognized path to value for labs
  • Thinking Machines emerged as an unexpected open-model leader, with their finetuning service generating hundreds of millions in annual revenue while releasing top U.S. open-weight models
  • Chinese labs maintain sustained momentum with new entrants like Xiaomi and releases from Tencent, Moonshot AI, and Meituan, complicating the open vs. closed model landscape
  • Licensing strategies are diverging: Apache 2.0 (Tencent, Poolside), OpenMDW (Poolside), and revenue-share/noncommercial licenses (Kimi K3) reflect competing visions for open model commercialization

Why It Matters

The failure of consolidation predictions to materialize signals a more competitive and fragmented AI ecosystem than anticipated, with open models carving out significant commercial and technical space. For practitioners, this means open-weight models are no longer second-class citizens—they are competitive, well-licensed, and commercially viable, reshaping deployment and fine-tuning strategies. The licensing diversity (permissive vs. restrictive) also introduces strategic considerations for enterprises choosing which models to adopt.

Technical Details

  • Inkling by Thinking Machines: A 975B-A41B multimodal MoE supporting text, image, and audio inputs with text output; a smaller 276B-A12B variant is also released and noted as highly competitive for its size. Positioned as a fine-tuning base via the Tinker commercial service.
  • Hy3 by Tencent: A 295B-A21B MoE improving over its predecessor across all metrics; notably switched from a restrictive custom license to Apache 2.0. Demonstrated ability to prove a 50-year-old math problem.
  • Laguna-S-2.1 by Poolside: An 118B-A8B MoE that fits on a single DGX Spark, newly pre- and post-trained. Released under the OpenMDW license (Apache 2.0-like with stronger AI-specific legal backing). Full evaluation trajectories published transparently.
  • DeepSeek-V4-Flash-0731 by DeepSeek-AI: Released one day after OpenAI cut its smallest model's prices by 80%; beats Luna at the pareto frontier in performance per parameter. The larger Pro variant was underwhelming in initial V4 releases.
  • Kimi-K3 by Moonshot AI: Described as the biggest open model release in some time, released under a noncommercial license requiring commercial agreements for inference and fine-tuning providers. Raises policy questions about U.S.-China AI business relationships.
  • LongCat-2.0 by Meituan: A 1.6T-parameter MoE from the Chinese "DoorDash"; strong on benchmarks but not the most capable for its size class beyond benchmark performance.

Industry Insight

  • The open-model ecosystem is maturing faster than expected, with companies like Thinking Machines proving that open-weight releases can generate hundreds of millions in revenue through finetuning services—challenging the assumption that only closed models are commercially viable.
  • Licensing is becoming a strategic differentiator: permissive licenses (Apache 2.0, OpenMDW) attract developer adoption, while restrictive licenses (Kimi K3) attempt to capture commercial value and navigate geopolitical risk, creating a fragmented landscape enterprises must navigate carefully.
  • Chinese labs continue to compete aggressively on both open and closed fronts, suggesting that U.S.-centric consolidation narratives may underestimate the global pace of open-model development and the role of revenue-share licensing as a middle ground.

TL;DR

  • 预测的AI实验室整合并未如期发生,更多公司仍在投入数亿至数十亿美元训练模型并开源发布
  • Token需求持续高涨,"构建token机器"成为AI价值新路径,开源微调服务可带来数亿美元年收入
  • 中国实验室(小米、美团等)保持强劲节奏,开源模型市场份额争夺进入关键时期
  • 许可证策略成为战略变量:Apache 2.0、OpenMDW与非商业许可证各有利弊,影响商业与政策格局
  • 多模态MoE架构成为主流,参数规模从百亿到万亿级并行发展,性价比与可部署性并重

为什么值得看

本文对开源模型生态进行了系统性梳理,揭示了整合预期与现实发展的背离,为从业者理解开源模型的商业化路径和政策风险提供了关键洞察。

技术解析

  • Inkling by Thinking Machines:975B-A41B多模态MoE模型,支持文本/图像/音频输入和文本输出,定位为微调基础模型;同时发布276B-A12B紧凑版本,在同类规模中竞争力突出,其微调服务年营收达数亿美元。
  • Hy3 by Tencent:295B-A21B MoE架构,全指标超越前代;关键变化是从限制性自定义许可证转向Apache 2.0开源协议,模型还成功证明了一道50年历史的数学问题。
  • Laguna-S-2.1 by Poolside:118B-A8B MoE模型,经预训练和后训练优化后可部署于单卡DGX Spark;采用OpenMDW许可证(类Apache 2.0但针对AI模型优化法律条款),并公开完整评估轨迹,透明度领先。
  • DeepSeek-V4-Flash:在OpenAI降价80%后次日发布更新,在帕累托前沿超越Luna模型;Flash版本在参数效率上表现优异,而Pro版本相对逊色。
  • Kimi K3 by Moonshot AI:近期最大规模开源模型之一,采用非商业许可证,要求推理和微调提供商签订商业协议;该策略引发政策讨论,可能为美国政府限制中美AI商业往来提供工具。
  • LongCat-2.0 by Meituan:1.6T参数MoE模型,虽在基准测试中非同规模最强,但展示了中国实验室持续投入大规模训练的能力。

行业启示

  • 整合预期落空表明AI基础设施投资回报路径比预想更分散,"token经济"和开源微调服务成为新盈利模式,创业者可关注垂直微调市场机会。
  • 许可证策略正成为地缘政治与商业博弈的交叉点,Apache 2.0等宽松协议与限制性商业许可证将塑造不同的生态格局,企业需评估合规与政策风险。
  • 中国实验室在开源模型领域持续发力且节奏稳定,美国开源生态需通过透明度(如Poolside的评估公开)和部署便利性(如单卡可运行)建立差异化竞争力。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Open Source 开源 LLM 大模型 Training 训练 Research 科学研究