Stanford HAI warns that averaging expert safety scores can erase good chatbot advice
StanfordHAI · x · 2026-07-25
Stanford HAI says AI developers often rely on mental health experts to judge whether chatbot answers are safe — but averaging expert scores can produce advice nobody would actually call good.
- The problem is not just disagreement, but the way disagreement gets collapsed into one number.
- The post points to a failure mode in safety evaluation for mental-health-style chatbot responses.
- It suggests that naive aggregation can hide meaningful differences in expert judgment.
More from Research
- NYU hires Simone Bombari to study memorization and privacy in large models — thegautamkamath · 2026-07-25
- Benchmark says frontier models still vary widely on antibody thermostability prediction — DeryaTR_ · 2026-07-25
- Kimi K3’s architecture is public, and the draft diagram shows KDA plus AttenRes — AccBalanced · 2026-07-25
- A Qwen3.6-27B merge blends reasoning and coding into one 27B model — pbaylies · 2026-07-25
- SIGIR 2026 recap says multivector retrieval works, now the focus is speed — CShorten30 · 2026-07-25
- A demo trains and visualizes RL policies inside tldraw, fully offline — max__drake · 2026-07-25