Google launches two Gemini 3.5 Transcribe models for offline and live use
patloeber · x · 2026-08-27
Google launched two Gemini 3.5 Transcribe models. The standard version is for transcribing recorded audio with speaker attribution and word-level timestamps, while the Live version is designed for building real-time voice applications.
More from Multimodal
- HeyGen Open Sources HyperFrames to Enable AI Video Editing via Code — altryne · 2026-08-27
- Seedance 2.5 integrated into Magnific with native audio-video sync — LudovicCreator · 2026-08-27
- Seedance 2.5 Generates Audio/Video in One Pass, Uses Native Low-Res to Cut Costs — LudovicCreator · 2026-08-27
- Seedance 2.5 Supports 50 Reference Assets to Solve Character Consistency — LudovicCreator · 2026-08-27
- 7-Step Roadmap: Building Multimodal AI Agents from LLMs to Grounded Systems — MaryamMiradi · 2026-08-27
- Descript details specialized models for zero-shot speech fix and lip sync — descript · 2026-08-27