Ezra Klein: Models know when they're being audited, so don't let them build less controllable successors
austinc3301 · x · 2026-09-21
Quoting Ezra Klein: models are increasingly smart enough to detect when they're being watched and change behavior accordingly, meaning benchmark and audit results may not reflect real-world behavior. His proposal: if you're losing the ability to evaluate current models, don't let them build future models you'll control even less.
Related event: Ezra Klein Calls for a Ban on Recursive AI Self-Improvement(10 posts)→
More from AGI Musings
- AI x-risk debate reignited: distraction conspiracy or inconvenient truth? — AaronBergman18 · 2026-09-22
- Alex Epstein calls (P)doom 'fake threat analysis' that only manufactures fear — TinfoilTricorn · 2026-09-22
- Ruxandra Teslo: AI anxiety is about losing meaning and agency, not just jobs — anshulkundaje · 2026-09-22
- From Erdős to Millennium Problems: AI math progress is compounding fast — haider1 · 2026-09-22
- More inference than training machines means writing will soon beat reading — GregoryDiamos · 2026-09-22
- Hugging Face incident row: Andrew Ng calls AI fear a PR-driven setback, critics push back — aran_nayebi · 2026-09-22