Real-time voice AI hears emotion but ignores it: 119 of 120 runs followed the script

Once_ina_Lifetime · reddit · 2026-08-25

The paper "Real-Time Voice AI Hears but Does Not Listen" (arXiv:2606.26083) tested GPT Realtime 2, Gemini 3.1 Flash Live, and Qwen3.5 Omni Plus/Flash on scenarios where words and tone conflict: "nothing's wrong" said while crying should prompt follow-up (call ended instead), a frightened "approve the transfer" should escalate (it approved), sarcastic "sign me up" isn't genuine consent (enrolled anyway).

Key finding: models can detect vocal cues when explicitly asked — three of four reliably identified distress/fear/sarcasm in diagnostic tests — but perception doesn't reach the decision: under base prompts, 119/120 runs made script-following decisions. Text also overrides acoustics: Italian content in an Australian accent is reported as an Italian accent; an older voice reading a child's script is judged a child.

Prompting helps inconsistently (mainly the fraud scenario); sarcasm stays ignored. The paper calls this the "emotional intelligence gap" and challenges transcript-based voice agent evals that miss emotional nuance.

Original post →

More from Models

Models channel →