OpenAI's CoT monitorability hit: paper authors double down on 'fragile' AI safety window

GaryMarcus · x · 2026-09-03

Following The Information's report that OpenAI's new techniques reduce chain-of-thought monitorability, Gary Marcus urged readers to revisit the 2025 paper "Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety" (co-authored by Yoshua Bengio, Beth Barnes, Neel Nanda and others). Its core claim: CoT monitoring is imperfect and lets some misbehavior slip through, but it's a promising oversight method worth investing in. Co-author Tomasz Korbak reaffirmed the paper's stance that monitorability is fragile yet preservable.

Original post →

More from Safety

Safety channel →