Stanford Study Exposes Major Flaw in AI Mental Health Safety Testing

StanfordHAI · x · 2026-07-30

As AI chatbots are increasingly used as counselors, developers rely on mental health experts to evaluate model responses for safety. However, new research from Stanford HAI highlights a critical flaw in this process: experts rarely agree on what constitutes "safe" advice.

In the study, three board-certified psychiatrists evaluated 360 AI responses to synthetic mental health prompts. The findings revealed significant disagreement among the experts' ratings, leaving AI developers struggling to find clear ways to improve mental health safety. The research suggests that traditional benchmark testing relying on single-expert scoring is unreliable and proposes a new path forward for evaluating AI safety.

Original post →

More from Safety

Safety channel →