Stanford study: API benchmark scores run 3.4 points higher than chatbot interfaces

StanfordAILab · x · 2026-09-17

A Stanford team (Jennifer Wang, Joachim Baumann, Daniel E. Ho, Sanmi Koyejo; EMNLP 2026) audited 7 systems across ChatGPT, Claude, and Gemini on 9 benchmarks and found API evaluations don't transfer to chatbot interfaces.

Key findings:

Implication: eval reports should state the access surface, and audits via API may certify a different system than the one most people actually use.

Related event: Stanford Study: API Audit Results Don't Transfer to Chat Interfaces(4 posts)→

Original post →

More from Safety

Safety channel →