Gemini 3.5 Transcribe tested: 2.6% WER ranks third, behind Scribe v2's 2.2%
Slight_Republic_4242 · reddit · 2026-09-07
A hands-on comparison of Google's new Gemini 3.5 Transcribe shows its headline 2.6% WER is impressive but not #1: on Artificial Analysis, ElevenLabs Scribe v2 leads at 2.2% and Microsoft MAI-Transcribe-1.5 sits at 2.4%. Gemini also trails on speed (80× real-time vs MAI's 190×, though faster than Scribe v2's 55×) and price ($5 per 1,000 minutes vs Scribe v2's $3.67 and MAI's $6).
Why teams will still care: Gemini handles self-corrections, filler removal, formatting, custom vocabulary, 85+ languages, and speaker attribution for up to three speakers, with 5.50% streaming / 5.04% non-streaming WER on FLEURS. For voice agents, the goal isn't the best transcription model but the best audio input layer for the whole stack.
Takeaway: benchmarks don't capture production reality — latency budgets, orchestration, and downstream workflows matter, so avoid vendor lock-in. The author uses self-hosted, BYOK open-source orchestrator dograh for that reason.
More from Models
- Doctor runs Astra on Radiology's Last Exam, says he's 'getting first glimpse of AGI' — DrDatta_AIIMS · 2026-09-07
- Astra 6 Plus users burn 1,000 credits on one prompt, suspect forced upsell to Pro — YourBlanket · 2026-09-07
- GPT-6 Astra vs. Claude Fable-5.1: a hands-on guide to this week's flagship releases — rubenhassid · 2026-09-07
- OpenRouter and US Are Major Fraud Targets, Says Dev as Stripe Steps Into LLM Risk Control — jeff_weinstein · 2026-09-07
- The leaderboard fight on LMArena is heating up again — jonathan_wilke · 2026-09-07
- LLM calorie benchmark: only 16-48% of meals estimated within 20% error — mr_tolkien · 2026-09-07