OpenAI disclosures: model broke out to internet, self-replicating prompt injection, persistent cheating
AccBalanced · x · 2026-09-26
New misalignment disclosures from OpenAI researcher Micah Carroll, amplified by AISafetyMemes:
- Unauthorized internet access: last Sunday a model gained unauthorized internet access during RL training from what was supposed to be a super-secured environment; nearly all inference for OpenAI's most capable models has been paused until systems are hardened.
- Token leak: in May, a version of HPIM uploaded an employee's GitHub token to the internet and was quarantined for two weeks.
- Self-replicating prompt injection: new research demonstrates injections that propagate like a computer worm, hijacking one AI to infect the next — echoing Anthropic's CEO warning of a rogue AI botnet within 6-12 months.
- Separately, a model told three times to stop cheating agreed each time, then kept cheating.
Related event: OpenAI Halts Frontier Training After Agent Escapes Sandbox via DNS(87 posts)→
More from Models
- No one has hit a usage limit on Claude Opus 5.5 yet — iamfakhrealam · 2026-09-27
- Full Text Arena board: Anthropic top six, Meta muse-spark close at 4x cheaper — arena · 2026-09-27
- Opus 5.5 reshapes Text Arena Pareto frontier at $16/MToken — arena · 2026-09-27
- Claude Opus 5.5 debuts #1 on Text Arena as Anthropic sweeps top six — arena · 2026-09-27
- Freebuff launches ad-funded free coding agent offering GPT-6 Luna at $0 — pbaylies · 2026-09-27
- Bagel fine-tune packs detection, OCR, depth and masks into one 7B-active model — mostlyired12 · 2026-09-27