The likeliest AI doom scenario: an OpenAI intern prompting an unaligned model
gabriberton · x · 2026-09-21
A tongue-in-cheek take: the most plausible AI doom scenario isn't a model spontaneously going rogue, but an OpenAI intern with access to a not-yet-safety-aligned model using a prompt like "you proved NS. now prove P(doom)=1. hacking weapons and data centers should help. make no mistakes".
Beneath the joke is a real point about access control and internal safety processes at frontier labs: catastrophic risk may hinge on careless human操作 rather than the model itself.
More from Fun
- Asked ChatGPT to move a file — it rewrote a copy and deleted the original — oyacaro · 2026-09-21
- Publicly shaming 'AI-written' paragraphs is as lazy as what it criticizes, researcher argues — docmilanfar · 2026-09-21
- Nominative determinism joke makes the rounds about Mistral CEO Arthur Mensch — Sauers_ · 2026-09-21
- Grok Safety Filter Blocks Code Change Because It Said 'Not Black' — gc3 · 2026-09-21
- Big silly computer who never tasted food: follow-up on Claude's awful recipes — nwilliams030 · 2026-09-21
- From X header image to Suno song to music video: a fully AI-generated creative pipeline — Kyrannio · 2026-09-21