Gemini 3.5 Transcribe supports real-time streaming and 85 languages
ammaar · x · 2026-08-27
Retweeting the introduction of Gemini 3.5 Transcribe. The speech-to-text model features smart transcription, function calling, lower WER, custom vocabulary, and multi-speaker identification. It supports over 85 languages and real-time streaming.
More from Multimodal
- HeyGen Open Sources HyperFrames to Enable AI Video Editing via Code — altryne · 2026-08-27
- Seedance 2.5 Generates Audio/Video in One Pass, Uses Native Low-Res to Cut Costs — LudovicCreator · 2026-08-27
- Seedance 2.5 Supports 50 Reference Assets to Solve Character Consistency — LudovicCreator · 2026-08-27
- 7-Step Roadmap: Building Multimodal AI Agents from LLMs to Grounded Systems — MaryamMiradi · 2026-08-27
- Descript details specialized models for zero-shot speech fix and lip sync — descript · 2026-08-27
- Runway integrates Meta's Muse image model, expanding multimodal capabilities — runwayml · 2026-08-27