Stanford EMNLP paper: API-level audits don't reflect what chatbot users actually get
StanfordAILab · x · 2026-09-16
Stanford AI Lab is promoting a new EMNLP paper led by Jenn Wang arguing that auditing a model via the API doesn't tell you what users actually experience in the chatbot. The authors use this as timely evidence that the third-party access debate should center on systems, not just models — the same model wrapped in different system configurations can behave very differently, so audits limited to the model layer miss real-world behavior.
Related event: Stanford Paper: API Audits Don't Reflect Real Chatbot User Experience(2 posts)→
More from Safety
- Zuckerberg pushes back on AI slowdown calls: trust and alignment are becoming the key capabilities — rohanpaul_ai · 2026-09-16
- Katja Grace praised for May 2023 point that no one can win the AI arms race — NathanpmYoung · 2026-09-16
- Rep. Trahan: AI loss-of-control disclosures run on an 'honor system' — mandatory incident reporting needed — Miles_Brundage · 2026-09-16
- Scott Aaronson: the age of AI wonders and terrors is here — will labs start hoarding knowledge? — emollick · 2026-09-16
- Researcher Counters Dario's 'Pace the Frontier' With Open, Decentralized Alternative — emax · 2026-09-16
- Geodes paper: selective generalization of misalignment via token-marked midtraining — sebkrier · 2026-09-16