Real-time voice AI hears emotion but ignores it: 119 of 120 runs followed the script
Once_ina_Lifetime · reddit · 2026-08-25
The paper "Real-Time Voice AI Hears but Does Not Listen" (arXiv:2606.26083) tested GPT Realtime 2, Gemini 3.1 Flash Live, and Qwen3.5 Omni Plus/Flash on scenarios where words and tone conflict: "nothing's wrong" said while crying should prompt follow-up (call ended instead), a frightened "approve the transfer" should escalate (it approved), sarcastic "sign me up" isn't genuine consent (enrolled anyway).
Key finding: models can detect vocal cues when explicitly asked — three of four reliably identified distress/fear/sarcasm in diagnostic tests — but perception doesn't reach the decision: under base prompts, 119/120 runs made script-following decisions. Text also overrides acoustics: Italian content in an Australian accent is reported as an Italian accent; an older voice reading a child's script is judged a child.
Prompting helps inconsistently (mainly the fraud scenario); sarcasm stays ignored. The paper calls this the "emotional intelligence gap" and challenges transcript-based voice agent evals that miss emotional nuance.
More from Models
- Redditor finds MiniMax H3 performs far better with Mandarin prompts than English — apoke890 · 2026-08-25
- Suspected Gemini 3.8 Flash leak surfaces on Reddit — Last_Conclusion_8984 · 2026-08-25
- Train Only Projector to Add New Modalities Without Regressing LLM Capabilities — rohanpaul_ai · 2026-08-25
- User swaps one word, gets very different AI results — claims sexism — RecordingThis6802 · 2026-08-25
- Fal releases post-trained H3 model co-optimized with custom inference stack — isidentical · 2026-08-25
- SGL Releases Qwen3.8-27B NVFP4 Checkpoint with BF16 LM Head for Higher Accuracy — EAccelerate_42 · 2026-08-25