VoxParity benchmark: only 11 of 23 voice agents act on what they hear, not just read
Bhavik Mangla · hf · 2026-10-01
VoxParity tests whether voice agents act on audio cues — a mayday under a radio check, a medical monitor beeping, a frightened whisper — rather than just the transcript. Across 183 scenarios from 14 sectors, the transcript stays fixed while the audio changes the correct tool call.
Key findings:
- Only 11 of 23 transcript-capable systems beat a words-only null test;
- Errors skew toward the words: when audio calls for protection, systems carry out the routine request more often (41%) than they over-react on clean calls (12%);
- Leading systems' misses cluster on cues they heard, and they overrule heard resignation or confusion far more often than acute alarm;
- Describing the voice and stating the rule each recover part of the gap, but emotion cues remain unsolved.
More from Models
- Yacine publicly churns off OpenAI, burning through all his astra tokens first — yacineMTB · 2026-10-01
- Kimi K2.6 was trained with the Tinker API, sparking calls for quirkier models — cHHillee · 2026-10-01
- Google Korea offers college students a free one-year Google AI Plus subscription — Ashamed_Photo2123 · 2026-10-01
- Astra 6 Ultrafast Mode: 300 Tokens/Sec Changes Agent Workflows, But Tools Are Now the Bottleneck — soumitrashukla9 · 2026-10-01
- Anthropic retires Claude Opus 3 but keeps it on API and gives it an essay column — repligate · 2026-10-01
- Bindu Reddy: Gemini Argon pricing is 5x cheaper than Astra, but benchmarks look too good — bindureddy · 2026-10-01