Gemini 3.5 Transcribe Hits 84x Realtime at $5/1,000 min, Ranks #5 on AA-WER
ArtificialAnlys · x · 2026-08-27
Google released Gemini 3.5 Transcribe, ranking #5 on Artificial Analysis's AA-WER benchmark for non-streaming transcription at 2.6% WER, alongside Gemini 3.5 Transcribe Live, which streams continuously via the Live API at 4.0% AA-WER with 0.40s latency after end of speech.
It processes 84 seconds of audio per second (84x realtime) at $5 per 1,000 minutes. Among top-five accuracy models it's faster than ElevenLabs Scribe v2 (55x, $3.67) but slower than Microsoft MAI-Transcribe-1.5 (191x) and Smallest AI Pulse Pro (275x, $4). The two offerings target pre-recorded audio (Interactions API) and live streaming (Live API) respectively.
Related event: Google Launches Gemini 3.5 Transcribe Speech-to-Text Model(22 posts)→
More from Models
- MiniMax H3 Max tops video leaderboards via fal's post-training — ArtificialAnlys · 2026-08-27
- GLM-5.3 Flash Review: GPT-5.6 Level Performance at Ultra-Low Cost — zainhas · 2026-08-27
- Qwen 3.8 and GLM 5.3 Flash open models released, available for Dell on-prem deployment — _akhaliq · 2026-08-27
- Fal secures new funding as H3 Max model achieves fastest speed and best quality — isidentical · 2026-08-27
- Local model choice in 2026: Qwopus 9B vs Qwen 27B — Ammargok · 2026-08-27
- Users report Qwen3.8 outputting garbage after prolonged use — trashacct383 · 2026-08-27