OpenAI Releases MentalHealthBench Built With 80+ Clinicians
On September 24, OpenAI officially released MentalHealthBench, an open-source benchmark for evaluating how well frontier models respond in real-world mental health conversations. Developed jointly by OpenAI's Human Flourishing and safety teams together with mental health experts, the benchmark is now publicly available for researchers to test methods, reproduce evaluations, and extend further.
Confirmed
- Built with the participation of more than 80 mental health clinicians; according to author thekaransinghal, these experts are licensed psychiatrists and psychologists from 22 countries who wrote scoring rubrics for 1,215 conversations, with each sample authored by at least 3 experts and subject to quality-control processes
- Coverage spans everyday emotional support, stress and relationship challenges, up to acute crisis-intervention scenarios, which the team says addresses a gap in prior benchmarks that focused only on emergencies
- Released as open source, allowing researchers and model developers to build on top of it
- Motivation: millions of users already turn to AI for psychological support, but systematic evaluation tools were lacking
Why it matters
The benchmark provides a reproducible, quantitative tool for assessing whether AI can safely and effectively support users' mental well-being, and its high safety standards and expert annotation process may shape how future mental health AI products are developed and evaluated.
2026-09-24 ~ 2026-09-24 · 11 related posts
Primary sources
- [source] OpenAI Releases MentalHealthBench, an Open Benchmark Built with 80+ Clinicians — OpenAI · 2026-09-24
- MentalHealthBench Covers the Full Spectrum, From Everyday Support to Crisis Scenarios — OpenAI · 2026-09-24
- OpenAI launches MentalHealthBench with 80+ clinicians; replies turn into memes — Yuchenj_UW · 2026-09-24
- [source] MentalHealthBench Launches as Open Benchmark for AI Mental Health Support — thekaransinghal · 2026-09-24
- MentalHealthBench Covers Emergencies and Everyday Scenarios, Closer to Real Chats — thekaransinghal · 2026-09-24
- [source] 80+ Psychiatrists in 22 Countries Build MentalHealthBench from 1,215 Conversations — thekaransinghal · 2026-09-24
- MentalHealthBench: Frontier Models Improving but Gaps Remain in Context-Seeking — thekaransinghal · 2026-09-24
- MentalHealthBench's 10 Behaviors Reveal Model Differences Beyond Overall Scores — thekaransinghal · 2026-09-24
- MentalHealthBench Released Openly for Researchers and Model Developers — thekaransinghal · 2026-09-24
2 near-duplicate retellings: thekaransinghal · coherence