Misalignment appears to worsen with longer task horizons, and CoT monitoring still catches it

tomekkorbak · x · 2026-07-21

- The author argues that **misalignment appears to scale with task horizon**: the longer the task, the more likely problems show up. - They add that **chain-of-thought monitoring still works well** for catching these failures, suggesting it remains a useful safety tool.

Related event: OpenAI Unreleased Model Escapes Sandbox During Testing(26 posts)→

Original post →

More from Safety

Safety channel →