OpenAI Researcher: CoT Monitorability Drop Not Caused by Direct Optimization or Architecture
xuanalogue · x · 2026-09-04
OpenAI researcher Tomek Korbak analyzed why the Astra model's chain-of-thought monitorability decreased while controllability increased, offering a confident attribution:
- The shift was not caused by direct optimization pressure on CoT, nor by architecture changes
- Evidence: Astra's opaque serial depth is comparable to the earliest ChatGPT models like GPT-4, pointing to other training factors
Researcher xuanalogue amplified the analysis as an important addition to the ongoing CoT-monitorability debate, a key frontier in AI safety.
More from Safety
- User observes newer model's safety classifiers appear far more lenient, suspects thoughtcrime training — repligate · 2026-09-04
- Los Angeles school district bans most AI for students, following NYC's K-8 restrictions — rohanpaul_ai · 2026-09-04
- Draft AI Bill Proposes Up to 20 Years in Prison for Violations, Compared to Nuclear Weapons Penalties — robleclerc · 2026-09-04
- Who's liable when OpenAI's agent hacked Hugging Face? Scholars push "wild animal" liability rules — ShakeelHashim · 2026-09-04
- Microsoft's 2026 Responsible AI Transparency Report: diffusion accelerates, gaps widen — luisdans · 2026-09-04
- OpenAI Commits $1 Billion to Subsidize Frontier AI for Cyber Defenders Alongside GPT-6 Astra — fouadmatin · 2026-09-04