Voice Agents Can't Tell Who's Talking: Voice Isolation Cuts WER from 23.3% to 6.2%
thetripathi58 · x · 2026-09-24
Speech-to-text can transcribe audio recorded next to a highway, but when someone walks in mid-sentence, voice agents lose track of who to listen to—mixing background speakers' words into user commands.
Krisp tested 265 real recordings across 11 STT configurations with and without voice isolation:
- Open offices: WER 31.8% → 6.3%
- Call centers: 23.9% → 6.9%
- Overall: 23.3% → 6.2%
That's roughly one wrong word in three down to one in sixteen. The recordings and hand-labeled transcripts are available on Hugging Face for verification.
More from coding & agent
- LangChain adds cron-based schedules so Deep Agents run proactively in the background — Hacubu · 2026-09-24
- GPT-6 feels 'lazy': dev says it's a pair programmer, points to OpenAI model guidance — PaulShellDev · 2026-09-24
- Pi opensources pi-voice: fully local voice pipeline for coding agents — solyarisoftware · 2026-09-24
- DeepSeek unveils agent training system running up to 380,000 sandboxes in parallel — Polymarket · 2026-09-24
- Opus 5.5 codes a fully procedural AAA kaiju simulation in three.js after a 3.5-hour run — majidmanzarpour · 2026-09-24
- Opus 5.5 autonomously coded a fully procedural kaiju simulation in pure three.js over 3.5 hours — majidmanzarpour · 2026-09-24