OpenAI's MentalHealthBench scores clinicians below most AI models — and that reveals a flaw
r0ck3t23 · x · 2026-09-24
OpenAI's MentalHealthBench (1,215 synthetic conversations, rubrics from 80+ clinicians across 22 countries) scored licensed clinicians at 38.5%, below most models — GPT-6 Astra hit 57.3%. The author argues the rubric rewards coverage over concise, clinically apt responses: writers given the rubric score 99%, and clinician and user rubric definitions only overlap 25.7% in weight. The grader is GPT-5.6 Sol and the top two models are OpenAI's own, though the benchmark is fully open and framed as a diagnostic. Context: APA's 2026 survey found 35% of psychologists say patients use AI as an additional mental health professional.
Related event: OpenAI Releases MentalHealthBench Built with 80+ Clinicians(13 posts)→
More from Models
- Opus 5.5 keeps trying to run rm -f; users must repeatedly beg it not to — gandamu_ml · 2026-09-24
- Third-party benchmark: AssemblyAI Universal 3.5 Pro tops 15 STT models at 1.93% WER and 489ms latency — AssemblyAI · 2026-09-24
- Why Jev may threaten frontier labs more than DeepSeek: an API so cheap everyone finds the waste — jobergum · 2026-09-24
- ChatGPT reportedly removes message cap on GPT-5.6 Luna for free users (unverified) — Aiden_Tech_Ai · 2026-09-24
- Grok 4.7 enters AutoResearchExam live leaderboard, ranks No.3 at 30min and No.4 after 24h auto-research — AlexGDimakis · 2026-09-24
- Researcher _xjdr: not liking astra, may go back to 5.6, eyeing Opus 5.5 and dsv4.1 flash — _xjdr · 2026-09-24