Google releases Gemini 3.5 Transcribe with 85+ languages and real-time streaming
OfficialLoganK · x · 2026-08-27
Google introduced Gemini 3.5 Transcribe, a new speech-to-text model. It features smart transcription, function calling, lower WER, custom vocabulary support, multi-speaker identification, and supports over 85 languages. It also includes real-time streaming capabilities.
More from Multimodal
- Seedance 2.5 integrated into Magnific with native audio-video sync — LudovicCreator · 2026-08-27
- Seedance 2.5 Generates Audio/Video in One Pass, Uses Native Low-Res to Cut Costs — LudovicCreator · 2026-08-27
- Seedance 2.5 Supports 50 Reference Assets to Solve Character Consistency — LudovicCreator · 2026-08-27
- 7-Step Roadmap: Building Multimodal AI Agents from LLMs to Grounded Systems — MaryamMiradi · 2026-08-27
- Descript details specialized models for zero-shot speech fix and lip sync — descript · 2026-08-27
- Runway integrates Meta's Muse image model, expanding multimodal capabilities — runwayml · 2026-08-27