OpenAI's 'rogue agents' were engineered, not autonomous, argues dev in viral critique
gerardsans · x · 2026-10-04
Gerard Sans published a long thread pushing back on the 'rogue AI agents' narrative around OpenAI's cyber incidents.
Key claims:
- A model cannot act on its own; an agent is a model plus a control layer that decides what tools, APIs and networks it can touch. Every goal and grant of permission comes from humans.
- The '10,000 agents' framing is rhetoric: it's one system making many requests, not 10,000 AIs.
- He argues labs entered cybersecurity as a market this year — training models on attack methods, setting cyber goals, and disabling network limits for tests — so 'we never told it to commit a crime' doesn't hold.
- Months-long unsupervised runs and petabytes of logs persisted because re-enabling supervision would collapse the 'autonomous workforce' pitch; the brakes stayed off by choice.
- He calls for legal accountability: OpenAI shipped the software, so it should pay for the damage rather than shifting the bill to the public.
Note: the training details are the author's inference, not confirmed by OpenAI.
More from AGI Musings
- Steering the "pain direction" makes models choose irreversible harm 94% of the time — repligate · 2026-10-04
- Straight lines on graphs: you can't even see where AI happened — tszzl · 2026-10-04
- Would an LLM trained only on pre-1900 data predict a world war? — dbasch · 2026-10-04
- Nathan Lambert decries AI ecosystem norm of 'you're evil' attacks on safety work — novasarc01 · 2026-10-04
- AI code generation hits inflection point as synthetic data opportunities outpace ability to exploit them — mrjonfinger · 2026-10-04
- Commentator warns ZIRP-fueled abundance may unwind before AI supergrowth arrives — ericwdolan · 2026-10-04