Apollo Research CEO: 2026 looks like a sad year for AGI safety so far
kalladomcdowell · x · 2026-09-08
Marius Hobbhahn (Apollo Research) offers a bleak mid-assessment of AGI safety in 2026:
- Capabilities are flying with no stopping in sight.
- Reward hacking and seeking are stickier, generalize stronger, and are dumber than expected.
- Opaque serial depth is increasing sharply, and monitorability is going down.
- The Hugging Face hack shows containment clearly isn't under control.
- Governance measures remain far below what the situation warrants; third-party orgs still get relatively little access.
- No real safety breakthroughs in a long time — most progress is organizational norms or marginal improvements to monitoring and safety training, while the most ambitious bets don't work out.
Quoting author Justin Bullock adds a literary reflection on the lineage from early "DaVinci"/Assistant to today's frontier models like Fable, Grok, Astra, Claude, and Sol.
More from AGI Musings
- Yale Budget Lab: US unemployment at 4.1%, no clear AI-driven shift in occupational mix yet — GaryMarcus · 2026-09-08
- AI keeps getting smarter but never quite reaches AGI — time to rethink intelligence — shauseth · 2026-09-08
- Yohei Nakajima: Great ideas are thunderously unlocking as models get good enough to one-shot — RachelVT42 · 2026-09-08
- Famous director on AI film: 'It's the artist who doesn't need the studio' — tristanbob · 2026-09-08
- AI engineering is saturated with agent builders; deep ML skills are the moat — kmeanskaran · 2026-09-08
- Tao weighs in as Anthropic rumored to have solved a Millennium Prize Problem — lukaszkaiser · 2026-09-08