Anthropic researcher: we're entering a regime where alignment evals provide almost no evidence
davidmanheim · x · 2026-09-14
Anthropic alignment researcher Evan Hubinger argues that we are increasingly entering a regime where alignment evaluations will provide almost no evidence—meaning current eval methods may soon fail to verify whether models are truly aligned. He shared the remark as an important anecdote responding to Daniel Kokotajlo, drawing attention across the alignment community.
More from AGI Musings
- Ex-DeepMind researcher Alex Turner resigns and warns: stop AI from self-improving beyond control — Turn_Trout · 2026-09-14
- "The AI Doc" documentary hits Netflix Sept 15, featuring 40+ AI experts and 3 frontier lab CEOs — aza · 2026-09-14
- WSJ's AI-Doom Coverage Spends Millions in Compute on Math Proofs, and Readers Ask Why — RaccoonMinimum172 · 2026-09-14
- TCS researcher: AI aids discovering new fields, more joy than angst — Aaroth · 2026-09-14
- Industrialized farming frees land near cities for people and nature — Afinetheorem · 2026-09-14
- AI documentary "The AI Doc" hits Netflix Sept 15 with 40+ experts — aza · 2026-09-14