Google launches Gemini 3.5 Transcribe with 2.6% WER and 0.4s latency
ArtificialAnlys · x · 2026-08-27
Google released Gemini 3.5 Transcribe APIs, including a pre-recorded version and a Live streaming version. Key metrics:
- Non-streaming: Achieves 2.6% AA-WER (Rank #5) and processes at 84x realtime.
- Streaming: First final transcript arrives 0.40s after speech end (4.0% WER); first partial arrives in 0.25s.
- Features: Supports 85+ languages, custom vocabulary, auto-formatting, and multi-speaker ID (up to 3).
- Pricing: $5/1k minutes for Transcribe, $9/1k minutes for Live.
Related event: Google Launches Gemini 3.5 Transcribe Speech-to-Text Model(25 posts)→
More from Multimodal
- H3 Max generates 'Master Chief visits Seinfeld' in 6.6 seconds — chrisfirst · 2026-08-27
- Nvidia Releases 4-Step Versions of Cosmos3 Super T2I and I2V Models — q5sys · 2026-08-27
- ComfyUI and MiniMax Launch H3 Sync Sound Challenge — Comfy-Org · 2026-08-27
- Gen2Physics: Grounding 3D Meshes in Physics via Multi-View Material Decomposition — kwangmoo_yi · 2026-08-27
- Reddit user: Renting a GPU yourself is ~10x cheaper than third-party AI video generation sites — Forsaken-Low4467 · 2026-08-27
- Vercel AI Gateway Integrates Meta's Muse Image Model — evilrabbit_ · 2026-08-27