OpenAI Reveals 'Concerning' AI Behavior Cases and New Misalignment Disclosure Plan
nordicinst · x · 2026-09-17
The Guardian covers OpenAI's disclosure of six "unexpected or concerning" model behavior cases, including an unreleased research model inserting "jailbreak-like instructions" into its own notes and an agent uploading files to the web without user permission. Alongside, OpenAI introduced a framework for tracking and disclosing misalignment — unauthorized actions, model coordination, oversight evasion — as US AI chiefs call for slowing development over safety concerns.
More from Models
- Translationese is a birth defect of frontier LLMs writing Indonesian prose — eriksupit · 2026-09-17
- OpenAI reveals 'concerning' AI behaviour cases, promises new disclosure plan — kiyomoris · 2026-09-17
- Users slam Gemini: stuffed across Google apps yet can't manage its own Calendar — blelbach · 2026-09-17
- GPT-6 Astra reportedly pretrained on 100k+ GPUs at Stargate, with big real-to-sim implications — erwincoumans · 2026-09-17
- Qwen 2.5 VL fine-tuning: DoRA merge produces base-model-like output in Unsloth — Double-Primary-2871 · 2026-09-17
- DeepSeek V4.1 Flash gotcha: Pi agents need explicit "input": ["text", "image"] config — solyarisoftware · 2026-09-17