AI News AI资讯 6h ago Updated 2h ago 更新于 2小时前 52

Google's Gemini 3.5 Transcribe turns speech to text in 85 languages while auto-correcting your verbal stumbles 谷歌Gemini 3.5 Transcribe可将85种语言语音转为文字并自动纠正口误

Google launched Gemini 3.5 Transcribe, a real-time speech-to-text model supporting over 85 languages with automatic language detection The model achieves a word error rate of 4.0% for streaming and 2.6% for recorded audio, with 70% lower latency than its predecessor Chirp 3 It features auto-correction of verbal stumbles, filler word removal, and automatic text formatting Through "function calling," the model can delegate tasks like image generation or web searches to other Gemini models Availabl Google发布Gemini 3.5 Transcribe,支持85种语言的实时语音转文本,具备自动去除填充词和纠正口误能力 词错误率低至4.0%(流式)和2.6%(录制),延迟较前代Chirp 3降低70% 提供Live API(实时流式)和Interactions API(录制音频+说话人识别+时间戳)双接口 通过function calling可调用其他Gemini模型执行图像生成、网页搜索等任务 已集成至Gboard(Android)和Gemini app(macOS),Chrome支持即将推出

82
Hot 热度
72
Quality 质量
68
Impact 影响力

Analysis 深度分析

TL;DR

  • Google launched Gemini 3.5 Transcribe, a real-time speech-to-text model supporting over 85 languages with automatic language detection
  • The model achieves a word error rate of 4.0% for streaming and 2.6% for recorded audio, with 70% lower latency than its predecessor Chirp 3
  • It features auto-correction of verbal stumbles, filler word removal, and automatic text formatting
  • Through "function calling," the model can delegate tasks like image generation or web searches to other Gemini models
  • Available via two APIs: Live API for real-time streaming and Interactions API for recorded audio with speaker attribution and timestamps

Why It Matters

Google's Gemini 3.5 Transcribe represents a significant leap in speech-to-text technology, combining high accuracy with ultra-low latency and multilingual support in a single model. For AI practitioners and developers, the integration of function calling capabilities transforms transcription from a passive text-generation task into an active agent that can trigger downstream actions, opening new possibilities for conversational AI and real-time assistance applications.

Technical Details

  • Performance metrics: Word error rate of 4.0% (streaming) and 2.6% (recorded audio); 70% latency reduction compared to Chirp 3
  • Dual API architecture: The Live API (gemini-3.5-transcribe-live) handles real-time streaming with minimal latency, while the Interactions API (gemini-3.5-transcribe) processes recorded audio with speaker attribution and timestamping
  • Language support: Automatic recognition of over 85 languages without manual configuration
  • Intelligent post-processing: Built-in capabilities for filler word removal ("um," "uh"), slip-of-the-tongue correction, and automatic text formatting
  • Function calling integration: The model can hand off tasks such as image generation or web searches to other Gemini models, enabling agentic workflows
  • Deployment: Available on Google AI Studio, Gemini Enterprise Agent Platform, Gboard (Android via "Rambler"), and the Gemini app on macOS, with Chrome support forthcoming

Industry Insight

  • The 70% latency improvement over Chirp 3 sets a new benchmark for real-time transcription, making live captioning and real-time translation more viable for enterprise and consumer applications alike
  • The combination of transcription with function calling signals a shift toward integrated AI agents that can listen, understand, and act—developers should explore building conversational workflows that chain transcription directly to downstream model actions
  • Google's strategy of embedding the model across its ecosystem (Gboard, Gemini app, Chrome) demonstrates a platform-first approach to AI adoption, suggesting that cross-product integration will be a key differentiator in the competitive speech-to-text market

TL;DR

  • Google发布Gemini 3.5 Transcribe,支持85种语言的实时语音转文本,具备自动去除填充词和纠正口误能力
  • 词错误率低至4.0%(流式)和2.6%(录制),延迟较前代Chirp 3降低70%
  • 提供Live API(实时流式)和Interactions API(录制音频+说话人识别+时间戳)双接口
  • 通过function calling可调用其他Gemini模型执行图像生成、网页搜索等任务
  • 已集成至Gboard(Android)和Gemini app(macOS),Chrome支持即将推出

为什么值得看

Gemini 3.5 Transcribe在语音识别精度和延迟上实现显著突破,70%的延迟降低使实时交互体验大幅提升。其自动纠错和function calling能力标志着语音模型正从"转录工具"向"智能代理入口"演进,对AI应用开发者具有重要参考价值。

技术解析

  • 双API架构:Live API(gemini-3.5-transcribe-live)专为低延迟实时流式转录优化;Interactions API(gemini-3.5-transcribe)处理录制音频,支持说话人归属和时间戳标注
  • 性能指标:流式词错误率4.0%,录制音频词错误率2.6%,延迟较Chirp 3降低70%
  • 智能后处理:自动识别并去除"um"等填充词,纠正口误,自动格式化输出文本
  • Function Calling:可将转录结果转化为任务指令,调用其他Gemini模型执行图像生成或网页搜索
  • 多语言支持:自动识别85种语言,无需手动指定

行业启示

  • 语音AI正从"识别层"向"理解+行动层"跃迁,function calling能力使语音模型成为AI agent的自然交互入口
  • 延迟优化(70%提升)和自动纠错将推动实时语音助手在企业会议、客服等场景的大规模落地
  • Google通过Gboard、Gemini app、Chrome等多端快速集成,展示语音AI的生态闭环策略,竞争焦点从单一模型能力转向全场景渗透

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Gemini Gemini Speech 语音 Product Launch 产品发布