Google announces Gemini 3.5 Transcribe for AI-powered speech-to-text
Google announced Gemini 3.5 Transcribe, an AI model that cleans up voice input by removing disfluencies like "ums" and "uhs" and editing self-corrections into polished text The model is approximately 70% faster than its predecessor Chirp 3, with live-speech error rates dropping from 7.32% to 5.5% It supports 85 languages and can handle up to three speakers in pre-recorded audio, while also respecting custom vocabulary for specialized jargon Already powering Gboard's "Rambler" feature on Pixel 11
Analysis
TL;DR
- Google announced Gemini 3.5 Transcribe, an AI model that cleans up voice input by removing disfluencies like "ums" and "uhs" and editing self-corrections into polished text
- The model is approximately 70% faster than its predecessor Chirp 3, with live-speech error rates dropping from 7.32% to 5.5%
- It supports 85 languages and can handle up to three speakers in pre-recorded audio, while also respecting custom vocabulary for specialized jargon
- Already powering Gboard's "Rambler" feature on Pixel 11 and the Gemini app on macOS, with broader rollout planned for Chrome and additional devices
- Developers can access the model through Antigravity, AI Studio, and the Gemini API starting immediately
Why It Matters
This represents Google's continued push to make voice input a seamless, production-ready interface for AI interaction, reducing the friction that has historically plagued speech-to-text workflows. For AI practitioners, it signals that disfluency removal and on-the-fly text editing are becoming table-stakes features rather than novelty additions, and the API availability opens the door for third-party integrations.
Technical Details
- Performance gains: ~70% faster voice-to-text latency compared to Chirp 3; live-speech error rate improved from 7.32% to 5.5%
- Disfluency handling: Automatically removes filler words ("ums," "uhs") and edits out self-corrections in real time to produce clean, coherent output
- Multilingual and multi-speaker support: Operates across 85 languages and can transcribe audio with up to three distinct speakers in pre-recorded files
- Custom vocabulary integration: Allows users to inject domain-specific terminology so the model preserves specialized jargon rather than normalizing it away
- Developer access: Available via Gemini API, AI Studio's build model, and Antigravity (which also provides screen context and chat history integration with user permission)
Industry Insight
- Voice-first input is rapidly maturing from a convenience feature into a core interaction paradigm; developers should prioritize integrating AI-enhanced transcription into any product that accepts spoken input
- The trade-off between accuracy and fidelity—where the AI rewrites rather than purely transcribes—means this tool is best suited for casual and productivity contexts, not legal or journalistic recording; product teams should make this distinction clear to users
- Google's strategy of layering this capability across its ecosystem (Gboard, macOS app, Chrome, API) suggests a deliberate play to lock in voice input as a default UX pattern, making API access a strategic move to extend that reach into third-party applications
Disclaimer: The above content is generated by AI and is for reference only.