AI News AI资讯 4h ago Updated 1h ago 更新于 1小时前 52

Google's new AI transcription edits out your 'ums' and 'ahs' 谷歌新AI转录功能自动删除"嗯""啊"等语气词

Google launched Gemini 3.5 Transcribe, a major advancement over Chirp 3 with improved multilingual performance and reduced word error rates The model supports 85+ languages, automatic jargon detection, customized vocabulary, and voice-based natural editing capabilities Features include speaker attribution for up to 3 speakers, word-level timestamps, automatic text formatting, and filler word removal Initially rolling out in English for macOS Gemini app and Android Rambler dictation in select reg Google推出Gemini 3.5 Transcribe,支持85种以上语言及专业术语自动识别,转录准确率较Chirp 3显著提升 新增语音自然编辑、自动格式化、填充词过滤功能,支持最多3人对话分离及词级时间戳标注 Gemini 3.5 Live/Live Experimental更新延期,仅Transcribe功能今日向macOS用户及Android部分国家开放 开发者可通过Gemini API(AI Studio/Antigravity)获取公开预览版,Chrome支持即将推出

68
Hot 热度
65
Quality 质量
60
Impact 影响力

Analysis 深度分析

TL;DR

  • Google launched Gemini 3.5 Transcribe, a major advancement over Chirp 3 with improved multilingual performance and reduced word error rates
  • The model supports 85+ languages, automatic jargon detection, customized vocabulary, and voice-based natural editing capabilities
  • Features include speaker attribution for up to 3 speakers, word-level timestamps, automatic text formatting, and filler word removal
  • Initially rolling out in English for macOS Gemini app and Android Rambler dictation in select regions, with Chrome support coming soon
  • Developers can access the model via public preview in Gemini API through AI Studio and Antigravity

Why It Matters

This represents Google's continued push to dominate the AI-powered audio transcription space with enterprise-grade features like specialized jargon handling and multi-speaker attribution. The integration of voice-based editing and customized vocabulary makes it particularly valuable for professionals who work with technical terminology, medical language, or industry-specific content.

Technical Details

  • Multilingual Support: Automatically detects and transcribes content in more than 85 languages with improved wording error rates compared to previous Chirp 3 model
  • Custom Vocabulary: Users can provide specialized terminology to prevent manual editing of unique spellings and domain-specific jargon
  • Speaker Attribution: Automatically identifies and attributes speech for up to three speakers in pre-recorded audio with word-level timestamps
  • Voice Editing: Natural voice-based editing capabilities allow users to modify transcriptions conversationally rather than through text interfaces
  • API Access: Available through Gemini API in public preview via AI Studio and Antigravity for developer integration

Industry Insight

Google is strategically positioning Gemini 3.5 Transcribe as an enterprise-ready solution that bridges the gap between consumer convenience and professional accuracy requirements. The delayed launch of 3.5 Live and 3.5 Live Experimental suggests Google is prioritizing transcription capabilities over real-time conversational features, indicating market demand for reliable audio-to-text solutions. AI practitioners should monitor this for potential integration into workflow automation, content creation pipelines, and accessibility applications.

TL;DR

  • Google推出Gemini 3.5 Transcribe,支持85种以上语言及专业术语自动识别,转录准确率较Chirp 3显著提升
  • 新增语音自然编辑、自动格式化、填充词过滤功能,支持最多3人对话分离及词级时间戳标注
  • Gemini 3.5 Live/Live Experimental更新延期,仅Transcribe功能今日向macOS用户及Android部分国家开放
  • 开发者可通过Gemini API(AI Studio/Antigravity)获取公开预览版,Chrome支持即将推出

为什么值得看

该更新标志着Google在多语言语音转录领域实现技术跃迁,专业术语自适应能力将直接降低医疗、法律等垂直场景的转录成本。语音编辑交互模式的引入为AI音频处理工具树立了新标准,开发者可借此探索实时语音工作流集成。

技术解析

  • 多语言架构:支持85+语言自动检测,针对专业术语库提供自定义词汇表训练,减少领域专有名词误识别
  • 语音处理管线:集成Chirp 3升级版声学模型,实现词级时间戳标注与说话人分离(最多3方),填充词过滤算法优化自然语言流畅度
  • 交互创新:首创语音指令编辑模式,用户可通过口语指令直接修改转录文本,结合自动格式化引擎输出结构化文档
  • 部署路径:macOS Gemini应用优先上线,Android端通过Rambler语音输入功能试点,API层通过AI Studio开放开发者预览

行业启示

  • 多语言语音AI正从通用转录向垂直领域专业化演进,术语自适应能力将成为企业级服务差异化竞争的关键指标
  • 语音编辑交互范式突破传统文本修改限制,预示下一代AI音频工具将深度融合自然语言指令与实时处理管线
  • 模型发布节奏调整(Live功能延期)反映大厂在复杂语音模型上的技术审慎,建议开发者优先接入已稳定的Transcribe API构建应用原型

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Gemini Gemini Speech 语音 Product Launch 产品发布 LLM 大模型