OpenAI's MentalHealthBench scores clinicians below most AI models — and that reveals a flaw

r0ck3t23 · x · 2026-09-24

OpenAI's MentalHealthBench (1,215 synthetic conversations, rubrics from 80+ clinicians across 22 countries) scored licensed clinicians at 38.5%, below most models — GPT-6 Astra hit 57.3%. The author argues the rubric rewards coverage over concise, clinically apt responses: writers given the rubric score 99%, and clinician and user rubric definitions only overlap 25.7% in weight. The grader is GPT-5.6 Sol and the top two models are OpenAI's own, though the benchmark is fully open and framed as a diagnostic. Context: APA's 2026 survey found 35% of psychologists say patients use AI as an additional mental health professional.

Related event: OpenAI Releases MentalHealthBench Built with 80+ Clinicians(13 posts)→

Original post →

More from Models

Models channel →