Google launches Gemini 3.5 Transcribe with support for 85+ languages and streaming
osanseviero · x · 2026-08-27
Google announced Gemini 3.5 Transcribe, a new speech-to-text model featuring smart transcription, function calling, custom vocabulary, and multi-speaker identification. It supports over 85 languages, offers both batch and real-time streaming audio support, and is now integrated across Google's products.
More from Multimodal
- HeyGen Open Sources HyperFrames to Enable AI Video Editing via Code — altryne · 2026-08-27
- Seedance 2.5 Generates Audio/Video in One Pass, Uses Native Low-Res to Cut Costs — LudovicCreator · 2026-08-27
- Seedance 2.5 Supports 50 Reference Assets to Solve Character Consistency — LudovicCreator · 2026-08-27
- 7-Step Roadmap: Building Multimodal AI Agents from LLMs to Grounded Systems — MaryamMiradi · 2026-08-27
- Descript details specialized models for zero-shot speech fix and lip sync — descript · 2026-08-27
- Runway integrates Meta's Muse image model, expanding multimodal capabilities — runwayml · 2026-08-27