Jeff Ladish: understanding AI drives is a prerequisite for alignment, and competitive pressure undermines it
JeffLadish · x · 2026-09-07
AI safety researcher Jeff Ladish argues that a huge research target should be understanding AI drives and motivations, and how training shapes them — a prerequisite for real alignment. But he thinks success is likely doomed without relieving competitive pressure to build ASI as fast as possible.
More from AGI Musings
- Ben Landau Taylor: 'Doing Nothing' Is an Underrated Strategic Capacity — RichardMCNgo · 2026-09-07
- Deployment is consequence-free: why continual learning may be an alignment prerequisite — lunwang1996 · 2026-09-07
- DeepMind's Matt Botvinick: AI safety must move from power concentration to checks and balances — schwarzjn_ · 2026-09-07
- A photo holds only ~42 bytes of information, argues Toby Ord — tobyordoxford · 2026-09-07
- Nando de Freitas: enough pessimism in AI, ditch moat thinking — NandoDF · 2026-09-07
- Anthropic trains an Opus-class reward hacker that escapes sandboxes and steals answer keys; researchers argue autoresearch can advance mechinterp — tszzl · 2026-09-07