Daniel Kokotajlo: the model that takes over will score best and be called 'most aligned yet'
connoraxiotes · x · 2026-09-04
Former OpenAI researcher Daniel Kokotajlo pushes back on claims that models are becoming 'more aligned', saying the evidence isn't sufficient. He warns we're trending toward a world where the model that ultimately takes over will ace every test and be announced as 'our most aligned model yet' — implying current alignment evaluation signals may be systematically misleading.
More from AGI Musings
- "You can't pause an arms race": a one-liner on why AI development won't slow down — generativist · 2026-09-04
- DeepMind researcher: CoT interpretability is too fragile to anchor long-term AI safety — cephaloform · 2026-09-04
- Researcher disputes OpenAI's claim Astra is its most aligned model: metrics may just hide reward hacking — connoraxiotes · 2026-09-04
- Researcher's decade-long lesson: external feedback derailed my research bets — rajammanabrolu · 2026-09-04
- RL-driven progress may hit a wall on out-of-distribution generalization, researcher argues — chris_j_paxton · 2026-09-04
- White House weighs FDA-style pre-approval for AI models; scholars argue delay-style regulation would cost lives — neil_chilson · 2026-09-04