EQ-Bench 4 adds 16-turn personas to test emotional intelligence in chatbots
OpenAIDevs · x · 2026-07-24
EQ-Bench 4 is out, a new multiturn benchmark for emotional intelligence and social interaction in chat.
- It uses 16-turn conversations with simulated personas that have competing traits, so there is no single best interaction policy.
- The benchmark is designed to test trust-building, adaptability, repair, and whether a model can infer user preferences from limited context.
- The author argues that models often look good on static EQ quizzes but still fail in situational judgment, becoming sycophantic, aloof, scripted, or overconfident.
- The post also includes leaderboard results and qualitative behavior summaries for several frontier models.
Related event: EQ-Bench 4 Launches to Test AI Emotional Intelligence(3 posts)→
More from Research
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Fruit fly brain as an LLM: connectome-driven language model demo goes live — ngxson · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11
- The Waymo effect: how AI is quietly making research less collaborative — JohnHammersley · 2026-09-11