Anthropic researcher: AI devs believe their tech could cause human extinction within years
xuanalogue · x · 2026-09-10
Anthropic researcher @saprmarks (writing personally) lays out the AI risk landscape, sparking a debate over interpretability-first research:
- AI developers largely believe their technology could cause human extinction, possibly within a few years — the more senior the employee, the greater the concern.
- They continue anyway due to commercial incentives plus a belief they're racing less responsible labs.
- Unlike traditional software, AIs can't be "programmed" to behave as intended; we only interpret and control them post hoc.
Responding to Jacob's thread, critics ask why the community didn't invest in explainable, controllable AI from the start instead of retrofitting control.
Related event: Anthropic Researcher Says AI Developers Believe Extinction Risk Is Real(2 posts)→
More from AGI Musings
- How Would AI Actually Kill Us? From Bioweapons to Gradual Handover of Control — RobbWiller · 2026-09-10
- Anthropic's risk statement slammed for self-praise over frank talk on AI risks — ShakeelHashim · 2026-09-10
- Should Anthropic researchers quit loudly? AI safety circle debates — JMannhart · 2026-09-10
- Quarter of AI researchers put human extinction risk above 25%, survey finds — KatjaGrace · 2026-09-10
- Beff Jezos: p(1984) far exceeds p(AI Doom), doomers are useful idiots — beffjezos · 2026-09-10
- Beff Jezos: AI regulation push aims to ban open source and nationalize Anthropic as a weapon factory — beffjezos · 2026-09-10