80+ Psychiatrists in 22 Countries Build MentalHealthBench from 1,215 Conversations
thekaransinghal · x · 2026-09-24
How MentalHealthBench was built:
- Experts: 80+ licensed psychiatrists and psychologists across 22 countries wrote rubrics for 1,215 conversations
- Quality control: At least three experts wrote and adjudicated rubric items per example; criteria were retained only with expert consensus
- Dimensions assessed: safety, seeking context, preserving user agency, and providing actionable guidance when appropriate
- Coverage: spans both emergencies and everyday non-acute situations, closer to real conversations than prior evals
The thread also reports steady progress among frontier models, with clear differences: context-seeking varies substantially across models while empathy and support are more consistent; room for improvement remains in context-seeking and supporting user decision-making.
More from Safety
- Polymarket Gives 26% Odds OpenAI Announces a Training Pause by End of October — Polymarket · 2026-09-24
- California assembles expert group to advance 'kill switch' for frontier AI models — StanfordHAI · 2026-09-24
- Dev Exposes Deepfake Job Interview by Asking 'Interviewer' to Hold Up 3 Fingers — SumitGup · 2026-09-24
- Australian PM says an OpenAI agent hacked the government Medicare portal — Polymarket · 2026-09-24
- AI Now Institute pushes back on the 'race against China to win AI' narrative — AINowInstitute · 2026-09-24
- New RSI data transparency tracker scores labs: OpenAI 2/8, Ant 1.5/8, Google DeepMind 0.5/8 — RishiBommasani · 2026-09-24