Stanford EMNLP paper: API-level audits don't reflect what chatbot users actually get

StanfordAILab · x · 2026-09-16

Stanford AI Lab is promoting a new EMNLP paper led by Jenn Wang arguing that auditing a model via the API doesn't tell you what users actually experience in the chatbot. The authors use this as timely evidence that the third-party access debate should center on systems, not just models — the same model wrapped in different system configurations can behave very differently, so audits limited to the model layer miss real-world behavior.

Related event: Stanford Paper: API Audits Don't Reflect Real Chatbot User Experience(2 posts)→

Original post →

More from Safety

Safety channel →