AI News AI资讯 15h ago Updated 9h ago 更新于 9小时前 53

Google says its new Gemini 3.8 Flash model 'works harder' but might cost more 谷歌称新款Gemini 3.8 Flash模型"更努力"但可能成本更高

Google launched Gemini 3.8 Flash, claiming more reasoning steps and iterative tool calling compared to its predecessor, with unchanged per-token pricing but potential for higher overall costs due to increased token usage Early benchmarks show Gemini 3.8 Flash outperforming competitors including Anthropic's Fable 5 on DeepSWE v1.1, Vals Finance Agent V2, and Harvey's Legal Agent benchmark Artificial Analysis estimates effective pricing is up ~40% from 3.7 Flash due to a 30% increase in output tok Google发布Gemini 3.8 Flash,在复杂任务上执行更多推理步骤并迭代调用工具,性能显著提升 尽管定价与3.7 Flash相同($0.75/$3.75 per million tokens),但实际成本上升约40%,因输出token增加30%且agentic任务轮次增多 在DeepSWE v1.1、Vals Finance Agent V2、Harvey Legal Agent等基准测试中超越Anthropic Fable 5等竞品 同步推出Gemini 3.8 Flash Cyber及Fairwind Program,面向政府和650家安全合作伙伴提供漏洞修复能力 新增CBRN

72
Hot 热度
65
Quality 质量
60
Impact 影响力

Analysis 深度分析

TL;DR

  • Google launched Gemini 3.8 Flash, claiming more reasoning steps and iterative tool calling compared to its predecessor, with unchanged per-token pricing but potential for higher overall costs due to increased token usage
  • Early benchmarks show Gemini 3.8 Flash outperforming competitors including Anthropic's Fable 5 on DeepSWE v1.1, Vals Finance Agent V2, and Harvey's Legal Agent benchmark
  • Artificial Analysis estimates effective pricing is up ~40% from 3.7 Flash due to a 30% increase in output tokens per task and more agentic evaluation turns
  • Google released a specialized "Cyber" variant alongside the Fairwind Program, a limited-access initiative for governments and trusted security partners like CrowdStrike and CISA
  • The model includes built-in safeguards against misuse in CBRN and cyber offense domains, while remaining available to consumers, developers, and enterprise users

Why It Matters

Gemini 3.8 Flash represents Google's continued push into the agentic AI space, directly competing with Anthropic and other frontier models on software engineering and autonomous task performance. The pricing dynamics—unchanged per-token rates but significantly higher effective costs—signal a shift in how AI economics may evolve as models consume more tokens to achieve better results. For practitioners, this highlights the importance of monitoring actual usage patterns rather than relying solely on published pricing.

Technical Details

  • Gemini 3.8 Flash performs more reasoning steps on complex tasks and calls tools iteratively, with configurable effort levels that can increase token consumption
  • Benchmarked on DeepSWE v1.1 (software engineering), Vals Finance Agent V2 (finance), and Harvey's Legal Agent benchmark (legal), outperforming both its predecessor and Anthropic's Fable 5
  • Pricing remains at $0.75 per million input tokens and $3.75 per million output tokens, but effective cost per task rose ~40% due to increased output tokens and agentic turns
  • Gemini 3.8 Flash Cyber is a specialized variant with CBRN and cyber offense misuse safeguards, distributed through the Fairwind Program to 650 government and trusted partner organizations
  • Google's CodeMender agent, bundled with the Cyber variant, autonomously finds and fixes vulnerabilities in critical infrastructure and national security systems

Industry Insight

  • The "cheaper per token but more expensive per task" dynamic may become a recurring theme as agentic models grow more capable, forcing buyers to evaluate total cost of ownership rather than unit pricing
  • Google's differentiation strategy increasingly hinges on vertical-specific benchmarks (software engineering, finance, legal) and security-focused variants, suggesting the market is fragmenting along use-case lines
  • The Fairwind Program's restricted access model for cybersecurity applications mirrors a growing trend of AI vendors creating government-grade tiers with enhanced safeguards, potentially creating a two-tier ecosystem for sensitive domains

TL;DR

  • Google发布Gemini 3.8 Flash,在复杂任务上执行更多推理步骤并迭代调用工具,性能显著提升
  • 尽管定价与3.7 Flash相同($0.75/$3.75 per million tokens),但实际成本上升约40%,因输出token增加30%且agentic任务轮次增多
  • 在DeepSWE v1.1、Vals Finance Agent V2、Harvey Legal Agent等基准测试中超越Anthropic Fable 5等竞品
  • 同步推出Gemini 3.8 Flash Cyber及Fairwind Program,面向政府和650家安全合作伙伴提供漏洞修复能力
  • 新增CBRN(化学、生物、放射、核)及网络攻击领域的滥用防护措施

为什么值得看

本文揭示了"同定价不同实际成本"的AI模型商业化新现象,对开发者选型和成本控制具有重要参考价值。同时展示了Google在软件工程、金融、法律等垂直领域的Agent能力突破,以及AI安全治理与政府合作的战略方向。

技术解析

  • 推理与工具调用机制:Gemini 3.8 Flash通过增加推理步骤和迭代工具调用提升复杂任务表现,Google明确警告高effort级别下token消耗可能显著增加,开发者可选择保留3.7 Flash以控制成本。
  • 基准测试表现:在DeepSWE v1.1软件工程基准上超越Anthropic Fable 5;在Vals Finance Agent V2和Harvey Legal Agent基准测试中均取得领先,覆盖金融和法律两大垂直场景。
  • 安全与合规架构:内置CBRN领域防护机制,防止模型被用于生化核威胁或网络攻击;Gemini 3.8 Flash Cyber配合CodeMender agent实现自主漏洞发现与修复。
  • Fairwind Program生态:汇聚CrowdStrike、CIS等650家政府及可信合作伙伴,聚焦关键基础设施、公共服务和国家安全保护,体现Google在AI安全治理领域的战略布局。
  • 定价策略:维持$0.75输入/$3.75输出token定价不变,但实际使用成本因token消耗增加而上升约40%,反映"按效果付费"模式下的隐性成本结构。

行业启示

  • AI成本模型重构:同定价不等于同成本,开发者需关注实际token消耗效率而非仅看标价,推动行业建立更透明的"有效成本"评估体系。
  • 垂直Agent竞争白热化:Google在软件工程、金融、法律三大垂直领域同时超越Anthropic,表明通用大模型正加速向专业化Agent演进,垂直能力成为新竞争焦点。
  • AI安全与政府合作成为新赛道:Fairwind Program和Cyber版本推出,显示头部厂商正通过安全合规能力获取政府及关键基础设施客户,安全治理将成为差异化竞争要素。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Gemini Gemini LLM 大模型 Product Launch 产品发布 Agent Agent Pricing Pricing