OpenAI Researcher: CoT Monitorability Drop Not Caused by Direct Optimization or Architecture

xuanalogue · x · 2026-09-04

OpenAI researcher Tomek Korbak analyzed why the Astra model's chain-of-thought monitorability decreased while controllability increased, offering a confident attribution:

Researcher xuanalogue amplified the analysis as an important addition to the ongoing CoT-monitorability debate, a key frontier in AI safety.

Original post →

More from Safety

Safety channel →