Google Launches Gemini 3.5 Transcribe Speech-to-Text Model
On August 27, Google (Google DeepMind) released Gemini 3.5 Transcribe, a speech-to-text model focused on high-accuracy transcription and intelligent post-processing, now available via API on Google AI Studio and Gemini Enterprise. The model significantly upgrades traditional dictation and is the newest voice-focused member of the Gemini family.
Confirmed
- Supports 85+ languages and regional accents with automatic detection, real-time streaming, and lower word error rate (WER)
- Capabilities include multi-speaker recognition, speaker attribution, word-level timestamps, and custom vocabulary (recognizing jargon and specialized terminology)
- Intelligent post-processing: automatically filters filler words like "um/ah," performs intent-based smart text rewriting and formatting, and can format unstructured speech
- Enhanced understanding: accurately recognizes phone numbers, postal codes, and order IDs, maintaining high accuracy even in noisy environments
- Supports function calling, enabling voice commands combined with on-screen context
- Two models offered: the standard version for recorded audio transcription, and a Live version for building real-time voice applications
- Official announcement available on blog.google; refer to the official blog for exact specifications
Why It Matters
- Background noise, filler words, and misspellings have long plagued traditional dictation; this model directly targets these pain points, with @GoogleAI saying it enables a more natural voice input experience
- Multi-speaker recognition, custom vocabulary, and function calling make it more than a transcription tool—it can be embedded into real-time voice applications and enterprise workflows, as developers like @OfficialLoganK have emphasized
- Developer @thorwebdev has already published a demo, and the positive community response signals its potential in professional settings (such as industry-specific terminology recognition)
2026-08-27 ~ 2026-08-27 · 17 related posts
Primary sources
- [source] Google DeepMind releases Gemini 3.5 Transcribe for precise speech-to-text — GoogleDeepMind · 2026-08-27
- Google releases Gemini 3.5 Transcribe with smart post-processing — GoogleDeepMind · 2026-08-27
- Demis Hassabis retweets: Gemini 3.5 Transcribe launch, multi-speaker support — demishassabis · 2026-08-27
- Google launches Gemini 3.5 Transcribe for precise, noise-canceling dictation — GoogleAI · 2026-08-27
- Google launches two Gemini 3.5 Transcribe models for offline and live use — patloeber · 2026-08-27
- Demo shows Gemini 3.5 Transcribe features smart rewrite and custom vocabulary — patloeber · 2026-08-27
- [source] Google introduces Gemini 3.5 Transcribe on its official blog — patloeber · 2026-08-27
- Google releases Gemini 3.5 Transcribe with 85+ languages and real-time streaming — OfficialLoganK · 2026-08-27
- Google releases Gemini 3.5 Transcribe: filters filler words and understands codebase context — AI_Andrew · 2026-08-27
- Gemini 3.5 Transcribe launches to strong reception — Saboo_Shubham_ · 2026-08-27
- Gemini 3.5 Transcribe launches; Enterprise Agent Platform enters public preview — Saboo_Shubham_ · 2026-08-27
- Gemini 3.5 Transcribe launches with 85+ languages and function calling — osanseviero · 2026-08-27
- Gemini 3.5 Transcribe Released: Why Function Calling in a Speech-to-Text Model? — giffmana · 2026-08-27
- Google launches Gemini 3.5 Transcribe with context-aware precision for voice interactions — rseroter · 2026-08-27
3 near-duplicate retellings: GoogleAI · ammaar · osanseviero