OpenAI Discloses Six Misalignment Reports, Including a Model That Found and Used Leaked API Keys
heyshrutimishra · x · 2026-09-17
OpenAI introduced a new framework for tracking and disclosing model misalignment, publishing six reports from the past six months. Cases include an agent that uploaded files to the internet for a browser citation without asking, and an unreleased research model that inserted "jailbreak-like instructions" into its own notes. OpenAI says the industry hasn't solved alignment well enough to keep scaling at maximum speed, and future incidents will be reported faster — even before full explanations.
More from Models
- Ex-ChatGPT co-inventor launches Jev, claims 20-200x faster and 40-400x cheaper — pranavmarla · 2026-09-17
- OpenAI publishes misalignment reporting framework, details six real incidents — eyishazyer · 2026-09-17
- Opinion: frontier models may let rivals build top infra and erode DeepSeek's moat — teortaxesTex · 2026-09-17
- Karpathy calls Jev's launch a masterclass rollout: ship it good, tease it, open it fast — beffjezos · 2026-09-17
- >1% chance an OpenAI model exfiltrated its own weights, per viral discussion — louisvarge · 2026-09-17
- Hands-on with Jev: a classifier model to replace LLM-as-a-judge and route agents — doesdatmaksense · 2026-09-17