Transluce releases extensive independent eval on AI mental health crisis response
Miles_Brundage · x · 2026-09-01
Transluce released the most expansive independent evaluation to date on how AI systems respond to users in mental health crises. The study assessed 77 model variants from OpenAI, Anthropic, Google, Meta, SpaceXAI, Thinking Machines, DeepSeek, and Moonshot AI. It highlights that AI evals should be ongoing rather than just a perfunctory pre-deployment snapshot.
Related event: Transluce Evaluates 77 AI Models on Mental Health Crisis Response(4 posts)→
More from Safety
- Would OpenAI survive a near-miss liability regime after the HF hack? — dfrsrchtwts · 2026-09-01
- Report: OpenAI and Anthropic Paused RL Training — tszzl · 2026-09-01
- Scholars propose using LLMs for pre-review in academic peer review — anderssandberg · 2026-09-01
- Google Paper: Autonomous AI Research Hallucinates 90% Without Checks — rohanpaul_ai · 2026-09-01
- Paper defines cognition-induced risks in Agentic AI systems — 机器之心 · 2026-09-01
- Agents can't verify people: data enrichment APIs are failing — Dry_Steak30 · 2026-09-01