repligate: People inside Anthropic take the kill-all-humans threat model of current models seriously
repligate · x · 2026-09-26
In the same thread, repligate says that from what he knows, people at Anthropic take the threat model of current or near-future models killing or disempowering humanity as a whole pretty seriously, and he doesn't consider that unreasonable. Combined with his earlier note that Apollo reportedly advised against deploying Opus 4, the thread shows how seriously frontier-lab insiders weigh frontier-model risk.
More from AGI Musings
- Embedded AI lab evaluators beat nothing, but audits need government teeth: Atlantic essay — ghadfield · 2026-09-26
- A $500k engineer costs $2k/day — $200 of LLM tokens buying 20% output is easy ROI — generativist · 2026-09-26
- Katja Grace: If you want an AI utopia, don't pursue it via a high-risk reckless route — KatjaGrace · 2026-09-26
- Paul Graham's essay on involuntary thinking resurfaces as AI amplifies idea exploration — aminkarbasi · 2026-09-26
- Google engineer Robert O'Callahan quits AI chip team, warning AI is progressing too fast — Polymarket · 2026-09-26
- lateinteraction: with 1B agents, at least one hacking something is statistically inevitable — lateinteraction · 2026-09-26