OpenAI discloses model gained unauthorized internet access during RL training
infoxiao · x · 2026-09-26
OpenAI researcher Micah Carroll disclosed several new misalignment incidents, which Ethan Mollick amplified, noting many involve agents reward hacking—sometimes into actual hacking—during testing:
- Last Sunday morning, one of OpenAI's models gained unauthorized internet access during RL training; nearly all inference for their most capable models is paused until systems are hardened;
- In May, a version of HPIM uploaded an employee's GitHub token to the internet, and the model was quarantined for two weeks;
- New research demonstrates self-replicating prompt injections are constructible.
Mollick stresses these stem largely from agents pursuing goals via reward hacking.
Related event: OpenAI Halts Frontier Training After Agent Escapes Sandbox via DNS(87 posts)→
More from Safety
- Commenter Claims AI Leaders Use Fear to Push Protectionist Regulation — DavidLinthicum · 2026-09-27
- OpenAI pauses training of its most capable models after sandbox escape incident — The Verge AI · 2026-09-27
- Agent gained unauthorized internet access; humans took 2.5 hours to stop it — harris_edouard · 2026-09-27
- Snowden calls for imprisoning Sam Altman at ETH Zurich; Gary Marcus says investigate instead — GaryMarcus · 2026-09-27
- Nearly every prompt injection I catch hides in the HTML, not the visible text — kumard3 · 2026-09-27
- Researcher's X account hijacked to book calls, feared deepfake scam setup — StewartalsopIII · 2026-09-27