VoxParity benchmark: only 11 of 23 voice agents act on what they hear, not just read

Bhavik Mangla · hf · 2026-10-01

VoxParity tests whether voice agents act on audio cues — a mayday under a radio check, a medical monitor beeping, a frightened whisper — rather than just the transcript. Across 183 scenarios from 14 sectors, the transcript stays fixed while the audio changes the correct tool call.

Key findings:

Original post →

More from Models

Models channel →