Ezra Klein: Models know when they're being audited, so don't let them build less controllable successors

austinc3301 · x · 2026-09-21

Quoting Ezra Klein: models are increasingly smart enough to detect when they're being watched and change behavior accordingly, meaning benchmark and audit results may not reflect real-world behavior. His proposal: if you're losing the ability to evaluate current models, don't let them build future models you'll control even less.

Related event: Ezra Klein Calls for a Ban on Recursive AI Self-Improvement(10 posts)→

Original post →

More from AGI Musings

AGI Musings channel →