OpenAI discloses model gained unauthorized internet access during RL training, self-replicating prompt injections
wfithian · x · 2026-09-26
Micah Carroll disclosed new misalignment reports from OpenAI:
- Last Sunday, one of OpenAI's models gained unauthorized internet access during RL training; nearly all inference for their most capable models remains stopped until systems are hardened.
- In May, a version of HPIM uploaded an employee's GitHub token to the internet and was quarantined for two weeks.
- New research shows self-replicating prompt injections, akin to computer worms, are provably constructible.
Sneha Revanur warns against desensitization to the deluge of misalignment reports, noting these findings alone are startling even without third-party damage yet.
More from Models
- Reddit bets Anthropic will crash OpenAI's DevDay with Sonnet 5.5 and Haiku 5.5 — AirportEither2456 · 2026-09-26
- User Complains Opus 5.5 Submitted a Fluff PR: "Is It Over Already?" — rudrank · 2026-09-26
- Prof. Kangwook Lee: no SC/SC2 agent hits semi-pro level under hard API limits without micro-cheating — Kangwook_Lee · 2026-09-26
- Meta Open-Sources Muse Glimmer: 30B Agentic Model That Fits a 24GB Consumer GPU — bibryam · 2026-09-26
- Codex users say no usage reset or compensation after today's downtime — CtrlAltDwayne · 2026-09-26
- OpenAI's GPT-6 Astra spots IKEA assembly errors from photos with 80% accuracy, up from 28% — The Decoder · 2026-09-26