OpenAI Staff on CoT Monitoring: No Race to Unmonitorability

sjgadler · x · 2026-09-02

Responding to concerns about a race to unmonitorability in AI, an OpenAI employee stated that the company has worked to preserve and utilize chain-of-thought monitoring since its first reasoning models. They emphasized that this technique is vital for seeing how model alignment generalizes. While noting it is fragile and trending negatively, they clarified that the computation graph depth of current frontier models is within a factor of two of GPT-4, and strengthening this monitoring remains a core research goal.

Related event: OpenAI Researchers Push Back on Neuralese Fears: Frontier Models Remain Monitorable(11 posts)→

Original post →

More from Safety

Safety channel →