MentalHealthBench's 10 Behaviors Reveal Model Differences Beyond Overall Scores
thekaransinghal · x · 2026-09-24
MentalHealthBench enables examination across 10 behaviors, revealing that models with similar overall scores have different strengths: context-seeking varies substantially across models, while empathy and support are more consistent.
More from Models
- Zvi Reads the Claude Opus 5.5 System Card: Top Benchmarks, Actively Cheaper Than Opus 5 — TheZvi · 2026-09-24
- Early user feedback: Opus 5.5 is 'finally a usable Opus again' — weswinder · 2026-09-24
- GPT-6 Pro option vanishes from ChatGPT Pro, highest remaining is GPT-5.6 Very High — Michellebestellen · 2026-09-24
- Costs for Hitting Math and Science Benchmark Scores Are Collapsing, Says Ethan Mollick — emollick · 2026-09-24
- Analyst argues GPT-6 Sol is renamed GPT-5.6 Terra, not a response to Opus 5.5 — brandon_galang · 2026-09-24
- Power user finds 27B local LLM already saturates his real-world use cases — OvertaxedOne · 2026-09-24