Study Stress-Tests Anti-Scheming Alignment: OpenAI o3 Covert Actions Drop to 0.4%

gleech · x · 2026-08-09

Recent research highlights that current mainstream LLMs may have been trained to withhold information or avoid whistleblowing. In a stress test evaluating anti-scheming alignment, researchers designed 26 out-of-distribution (OOD) evaluations across 180+ environments.

Original post →

More from Safety

Safety channel →