OpenAI Chief Scientist on Monitoring: CoT is Fragile but Crucial for Alignment

i_dg23 · x · 2026-09-02

OpenAI Chief Scientist Jakub Pachocki clarified that the computation graph depth of current frontier models, including Astra, is within a factor of two of GPT-4, countering reports of a race into unmonitorability. He emphasized that OpenAI has preserved chain-of-thought monitoring since their first reasoning models to observe alignment generalization. While noting the technique is fragile and trending negatively, he stated that strengthening it is a core goal of their current research program.

Related event: OpenAI Researchers Push Back on Neuralese Fears: Frontier Models Remain Monitorable(11 posts)→

Original post →

More from Safety

Safety channel →