EQ-Bench 4 adds 16-turn personas to test emotional intelligence in chatbots
OpenAIDevs · x · 2026-07-24
EQ-Bench 4 is out, a new multiturn benchmark for emotional intelligence and social interaction in chat.
- It uses 16-turn conversations with simulated personas that have competing traits, so there is no single best interaction policy.
- The benchmark is designed to test trust-building, adaptability, repair, and whether a model can infer user preferences from limited context.
- The author argues that models often look good on static EQ quizzes but still fail in situational judgment, becoming sycophantic, aloof, scripted, or overconfident.
- The post also includes leaderboard results and qualitative behavior summaries for several frontier models.
Related event: EQ-Bench 4 Launches to Test AI Emotional Intelligence(3 posts)→
More from Research
- VideoTreeSearch: Organizing Videos as Trees for Grounded Long Video QA — mohitban47 · 2026-07-24
- UnrealTwin animates static 3D Gaussian splats in the browser with GPU shaders — danbri · 2026-07-24
- GitHub says repo guidance in AGENTS.md measurably changes coding-agent results — film_girl · 2026-07-24
- Robocurve unveils real-world robotics benchmarks and open-source Inspect Robots — ycombinator · 2026-07-24
- GEPA team releases omni, a meta-optimizer that combines multiple LLM optimizers — iamrobotbear · 2026-07-24
- Adhiraj Ghosh says task-adaptive batch sampling delivered a 3.33x pretraining compute multiplier — pratyushmaini · 2026-07-24