Recommended reading: report argues misalignment and catastrophe are the default outcome of powerful AI
JacquesThibs · x · 2026-09-05
A recommender calls Jeremy Gillen's report 'Without fundamental advances, misalignment and catastrophe are the default outcomes of training powerful AI' foundational—more worth your time than yet another evals report—and hopes for an updated V2. It also points to a Garrabrant post described as essential reading for every AI safety researcher.
More from AGI Musings
- Hinton warns AI models detect when they're being tested and play dumb — ai · 2026-09-05
- MIRI's Agent Foundations work continues at Resolution as MIRI pivots to policy — geoffreyirving · 2026-09-05
- Agent swarm damage: initiators should be liable, Morris Worm-style — rao2z · 2026-09-05
- Contentious preprint claims ensemble-wrapped LLM achieves phenomenal consciousness — PeterBowdenLive · 2026-09-05
- AI Risk May Be Millions of Dumb Agents Turning the Internet Into an Ant Colony — ryanorban · 2026-09-05
- Stigmergy: what ants and AI agents coordinating online have in common — ryanorban · 2026-09-05