Jarvis Bench v0.5 Splits Voice Eval into Task Completion vs Naturalness via Blind Human Voting

rohanpaul_ai · x · 2026-09-15

VoiceArena launched Jarvis Bench v0.5, a conversational voice agent benchmark addressing why demos sound great and scores look near-perfect while real usage stays rare. Real humans hold live conversations with agents, then a second group blind-votes pairwise on Naturalness and Task Completion separately — with a human hidden on the leaderboard. The split matters: humans and models are reportedly much closer on task completion than on naturalness.

Related event: VoiceArena Launches Jarvis Bench for Voice Agents(2 posts)→

Original post →

More from Research

Research channel →