Casper on METR/Redwood's OpenAI report: OpenAI controlled scope, negligence questions left out

StephenLCasper · x · 2026-09-23

Stephen Casper critiques the METR/Redwood embedded-evaluation report on OpenAI: it establishes a useful model and case study, but embedded evaluation spans a broad spectrum — from 'safety washing' to rigorous oversight with auditor access and whistleblowing. OpenAI controlled everything in scope and excluded the most important questions about negligence.

Related event: Researcher criticizes embedded evaluations as weak pacing substitute(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →