AI News AI资讯 1h ago Updated 1h ago 更新于 1小时前 48

Google announces Gemini 3.5 Transcribe for AI-powered speech-to-text 谷歌发布Gemini 3.5 Transcribe用于AI驱动语音转文字

Google announced Gemini 3.5 Transcribe, an AI model that cleans up voice input by removing disfluencies like "ums" and "uhs" and editing self-corrections into polished text The model is approximately 70% faster than its predecessor Chirp 3, with live-speech error rates dropping from 7.32% to 5.5% It supports 85 languages and can handle up to three speakers in pre-recorded audio, while also respecting custom vocabulary for specialized jargon Already powering Gboard's "Rambler" feature on Pixel 11 Google发布Gemini 3.5 Transcribe语音转文字模型,可自动去除口语中的"ums"和纠正内容,输出精炼文本 相比前代Chirp 3,语音到文本速度提升70%,实时语音错误率从7.32%降至5.5% 支持85种语言,最多3个说话者,可识别用户自定义专业词汇 已集成至Gboard Rambler、Gemini macOS应用、Antigravity、AI Studio及Gemini API 计划扩展至Chrome浏览器,支持网页任意文本框的语音输入

72
Hot 热度
65
Quality 质量
68
Impact 影响力

Analysis 深度分析

TL;DR

  • Google announced Gemini 3.5 Transcribe, an AI model that cleans up voice input by removing disfluencies like "ums" and "uhs" and editing self-corrections into polished text
  • The model is approximately 70% faster than its predecessor Chirp 3, with live-speech error rates dropping from 7.32% to 5.5%
  • It supports 85 languages and can handle up to three speakers in pre-recorded audio, while also respecting custom vocabulary for specialized jargon
  • Already powering Gboard's "Rambler" feature on Pixel 11 and the Gemini app on macOS, with broader rollout planned for Chrome and additional devices
  • Developers can access the model through Antigravity, AI Studio, and the Gemini API starting immediately

Why It Matters

This represents Google's continued push to make voice input a seamless, production-ready interface for AI interaction, reducing the friction that has historically plagued speech-to-text workflows. For AI practitioners, it signals that disfluency removal and on-the-fly text editing are becoming table-stakes features rather than novelty additions, and the API availability opens the door for third-party integrations.

Technical Details

  • Performance gains: ~70% faster voice-to-text latency compared to Chirp 3; live-speech error rate improved from 7.32% to 5.5%
  • Disfluency handling: Automatically removes filler words ("ums," "uhs") and edits out self-corrections in real time to produce clean, coherent output
  • Multilingual and multi-speaker support: Operates across 85 languages and can transcribe audio with up to three distinct speakers in pre-recorded files
  • Custom vocabulary integration: Allows users to inject domain-specific terminology so the model preserves specialized jargon rather than normalizing it away
  • Developer access: Available via Gemini API, AI Studio's build model, and Antigravity (which also provides screen context and chat history integration with user permission)

Industry Insight

  • Voice-first input is rapidly maturing from a convenience feature into a core interaction paradigm; developers should prioritize integrating AI-enhanced transcription into any product that accepts spoken input
  • The trade-off between accuracy and fidelity—where the AI rewrites rather than purely transcribes—means this tool is best suited for casual and productivity contexts, not legal or journalistic recording; product teams should make this distinction clear to users
  • Google's strategy of layering this capability across its ecosystem (Gboard, macOS app, Chrome, API) suggests a deliberate play to lock in voice input as a default UX pattern, making API access a strategic move to extend that reach into third-party applications

TL;DR

  • Google发布Gemini 3.5 Transcribe语音转文字模型,可自动去除口语中的"ums"和纠正内容,输出精炼文本
  • 相比前代Chirp 3,语音到文本速度提升70%,实时语音错误率从7.32%降至5.5%
  • 支持85种语言,最多3个说话者,可识别用户自定义专业词汇
  • 已集成至Gboard Rambler、Gemini macOS应用、Antigravity、AI Studio及Gemini API
  • 计划扩展至Chrome浏览器,支持网页任意文本框的语音输入

为什么值得看

Gemini 3.5 Transcribe代表了语音输入技术的重大进步,通过AI智能编辑功能显著提升语音转文字的准确性和用户体验。对AI从业者和开发者而言,这展示了语音交互领域的最新技术方向,也为多语言语音处理提供了新的参考方案。

技术解析

Gemini 3.5 Transcribe相比Chirp 3在速度上提升约70%,实时语音错误率从7.32%降至5.5%,虽然提升幅度有限,但对语音输入体验有实质性改善。模型支持85种语言和最多3个说话者的预录音频场景,并能通过自定义词汇表处理专业术语。

该模型具备智能语音编辑能力,可实时去除"ums"、"uhs"等口语填充词,并在用户自我纠正时动态编辑文本。技术架构已集成至Gemini API、AI Studio的build模型以及Antigravity应用,开发者可通过API直接调用。

行业启示

Google正将语音交互能力深度整合至其全平台生态,从Pixel设备到macOS应用再到Chrome浏览器,体现了语音输入作为主流交互方式的发展趋势。对开发者而言,Gemini API的开放意味着可以构建更智能的语音应用,而Chrome浏览器的集成预示着网页端语音交互体验将迎来重大升级。语音转文字技术的竞争焦点已从单纯的识别准确率转向语义理解和智能编辑能力。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Gemini Gemini Speech 语音 Product Launch 产品发布 LLM 大模型