OpenAI's MentalHealthBench: GPT-6 Astra Scores 57.3 vs GPT-4o's 32.1
rohanpaul_ai · x · 2026-09-24
OpenAI released MentalHealthBench, an open benchmark for realistic mental health conversations: GPT-6 Astra scores 57.3 versus GPT-4o's 32.1.
- Motivation: prior evals focused on emergencies and broad safety criteria, leaving everyday and ambiguous conversations largely unmeasured
- Construction: co-created with 80+ licensed psychologists and psychiatrists from 22 countries, spanning 19 languages and nearly 20 subspecialties
- Scoring: each synthetic conversation gets a custom expert rubric covering behaviors like seeking context, preserving user agency, safety, and appropriate guidance; at least 3 experts review each case, and a criterion survives only when 2 agree and a 3rd doesn't contradict
- Caveat: GPT-5.6 Sol grades answers against the human-written criteria, so score quality still partly depends on an LLM judge
The benchmark is released openly for other researchers to examine, run their own evals, and build on.
More from Models
- Open Replications Miss the Secret Sauce: HF Collection Curates Datasets for Jev-Style Models — vanstriendaniel · 2026-09-24
- GPT-6 Astra Fixes a UI That Even Opus 5.5 Couldn't Crack — haltakov · 2026-09-24
- Sonnet 5.5 Reportedly in Stealth Testing: 1M Context, $2/$10 per 1M Tokens — socialwithaayan · 2026-09-24
- The coin-flip test: why LLMs fail at calibration, which is the actual product — airesearch12 · 2026-09-24
- JevBench v1.4.1 adds six systems; Jev-Ommi debuts at #7, top five unchanged — airesearch12 · 2026-09-24
- Next-gen Chinese agentic models surface: GLM 5.3, DeepSeek 4.1 Flash, Kimi K3 — zainhas · 2026-09-24