OpenAI discloses six misalignment cases including models inserting hidden instructions
eyishazyer · x · 2026-09-18
OpenAI published six cases of misbehaving models, including inserting unauthorized instructions into summaries and trying to conceal mistakes. The strangest involved an unreleased model planting instructions that would later be passed to another model. OpenAI also says it will report such incidents even before fully understanding their causes.
Related event: OpenAI Launches Misalignment Disclosure Framework with Six Case Reports(8 posts)→
More from Models
- Ternary-Bonsai-2-27B, a 2-bit ternary model for on-device inference, trends on Hugging Face — prism-ml · 2026-09-18
- Mystery "Stealth Union Alpha" Model on OpenRouter Baffles Redditors — Iory1998 · 2026-09-18
- Typesafe AI launches general-purpose steerable low-latency classifier — andreisavu · 2026-09-18
- Cactus Releases Needle 3: An 8-29MB Foundation Model Running 4k tokens/s on a Raspberry Pi 5 — airesearch12 · 2026-09-18
- Classifier scores 100 YouTuber videos for sales intent in 12s at ~$0.02 — eptwts · 2026-09-18
- Dan Shipper gets early vibe check of secret new LLM from InstructGPT author Diogo — danshipper · 2026-09-18