Stanford study: API benchmark scores run 3.4 points higher than chatbot interfaces
StanfordAILab · x · 2026-09-17
A Stanford team (Jennifer Wang, Joachim Baumann, Daniel E. Ho, Sanmi Koyejo; EMNLP 2026) audited 7 systems across ChatGPT, Claude, and Gemini on 9 benchmarks and found API evaluations don't transfer to chatbot interfaces.
Key findings:
- API evaluations score 3.4 points higher in accuracy and 2.1 points higher in test-retest agreement than interface evaluations.
- For ChatGPT, the API-vs-interface gap rivals the API-only gap between GPT 5.3 and GPT 5.4 — switching access surfaces can degrade performance as much as a full model generation downgrade.
- Varying system prompts, sampling parameters, and reasoning settings failed to reliably reproduce interface behavior; whatever sits between endpoint and deployed product can't be reconstructed by auditors.
Implication: eval reports should state the access surface, and audits via API may certify a different system than the one most people actually use.
Related event: Stanford Study: API Audit Results Don't Transfer to Chat Interfaces(4 posts)→
More from Safety
- Anthropic and OpenAI pledge to embed third-party safety evaluators, but true independence remains an open question — RebeccaBellan · 2026-09-17
- Baseten partners with Goodfire and ba labs on safety and interpretability for open models — JJitsev · 2026-09-17
- Curtis Yarvin: The 'Rogue Agent' Panic Is Negligence Misread as Machine Rebellion — ivan_bezdomny · 2026-09-17
- WIRED: Washington won't regulate AI anytime soon as White House opposes oversight — nordicinst · 2026-09-17
- AI safety movement should learn from environmentalism: offer compelling alternatives before restriction — willcb · 2026-09-17
- Mel Mitchell: HF hackathon agents' 'loyalty' was a product of cooperative RL training — dhadfieldmenell · 2026-09-17