Google Launches Gemini 3.5 Transcribe Speech-to-Text Model
On August 27, Google (Google DeepMind) released Gemini 3.5 Transcribe, a speech-to-text model focused on high-accuracy transcription and intelligent post-processing, now available via API on Google AI Studio and Gemini Enterprise. The model significantly upgrades traditional dictation and is the newest voice-focused member of the Gemini family.
Confirmed
- Supports 85+ languages and regional accents with automatic detection, real-time streaming, and lower word error rate (WER)
- Capabilities include multi-speaker recognition, speaker attribution, word-level timestamps, and custom vocabulary (recognizing jargon and specialized terminology)
- Intelligent post-processing: automatically filters filler words like "um/ah," performs intent-based smart text rewriting and formatting, and can format unstructured speech
- Enhanced understanding: accurately recognizes phone numbers, postal codes, and order IDs, maintaining high accuracy even in noisy environments
- Supports function calling, enabling voice commands combined with on-screen context
- Two models offered: the standard version for recorded audio transcription, and a Live version for building real-time voice applications
- Official announcement available on blog.google; refer to the official blog for exact specifications
Why It Matters
- Background noise, filler words, and misspellings have long plagued traditional dictation; this model directly targets these pain points, with @GoogleAI saying it enables a more natural voice input experience
- Multi-speaker recognition, custom vocabulary, and function calling make it more than a transcription tool—it can be embedded into real-time voice applications and enterprise workflows, as developers like @OfficialLoganK have emphasized
- Developer @thorwebdev has already published a demo, and the positive community response signals its potential in professional settings (such as industry-specific terminology recognition)
2026-08-27 ~ 2026-08-27 · 25 related posts
Primary sources
- Google DeepMind releases Gemini 3.5 Transcribe for precise speech-to-text — GoogleDeepMind · 2026-08-27
- Google releases Gemini 3.5 Transcribe with smart post-processing — GoogleDeepMind · 2026-08-27
- Hands-on: Gemini 3.5 Transcribe achieves 2.6% WER, excellent post-processing — _philschmid · 2026-08-27
- Demis Hassabis retweets: Gemini 3.5 Transcribe launch, multi-speaker support — demishassabis · 2026-08-27
- Google launches Gemini 3.5 Transcribe for precise, noise-canceling dictation — GoogleAI · 2026-08-27
- Google launches two Gemini 3.5 Transcribe models for offline and live use — patloeber · 2026-08-27
- Demo shows Gemini 3.5 Transcribe features smart rewrite and custom vocabulary — patloeber · 2026-08-27
- Google introduces Gemini 3.5 Transcribe on its official blog — patloeber · 2026-08-27
- Google releases Gemini 3.5 Transcribe: filters filler words and understands codebase context — AI_Andrew · 2026-08-27
- Gemini 3.5 Transcribe launches to strong reception — Saboo_Shubham_ · 2026-08-27
- Gemini 3.5 Transcribe launches; Enterprise Agent Platform enters public preview — Saboo_Shubham_ · 2026-08-27
- Gemini 3.5 Transcribe launches with 85+ languages and function calling — osanseviero · 2026-08-27
- Gemini 3.5 Transcribe Released: Why Function Calling in a Speech-to-Text Model? — giffmana · 2026-08-27
- Google launches Gemini 3.5 Transcribe with context-aware precision for voice interactions — rseroter · 2026-08-27
- Hands-on: Gemini Transcribe 3.5 offers verbatim and smart modes — stefanjblos · 2026-08-27
- [source] Google launches Gemini 3.5 Transcribe with 2.6% WER and 0.4s latency — ArtificialAnlys · 2026-08-27
- Gemini 3.5 Transcribe Live Beats GPT Live Transcribe: 5.8% WER at 0.25s First Partial — ArtificialAnlys · 2026-08-27
- Gemini 3.5 Transcribe Live pricing higher than competitors at $9/1k mins — ArtificialAnlys · 2026-08-27
- Gemini 3.5 Transcribe Live costs ~$9/1k minutes, pricier than Cartesia and Deepgram — ArtificialAnlys · 2026-08-27
- Google launches Gemini 3.5 Transcribe for speech-to-text — ocean_protocol · 2026-08-27
5 near-duplicate retellings: GoogleAI · OfficialLoganK · ammaar · osanseviero · ArtificialAnlys