Researcher criticizes embedded evaluations as weak pacing substitute

Stephen Casper criticized the METR/Redwood embedded evaluation report on OpenAI, arguing that 'more evaluations' is a weak substitute for stricter pacing measures like hardware controls, and that OpenAI shaped the evaluation scope, leaving key oversights unaddressed.

2026-09-23 ~ 2026-09-23 · 2 related posts