OpenAI's CoT monitorability hit: paper authors double down on 'fragile' AI safety window
GaryMarcus · x · 2026-09-03
Following The Information's report that OpenAI's new techniques reduce chain-of-thought monitorability, Gary Marcus urged readers to revisit the 2025 paper "Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety" (co-authored by Yoshua Bengio, Beth Barnes, Neel Nanda and others). Its core claim: CoT monitoring is imperfect and lets some misbehavior slip through, but it's a promising oversight method worth investing in. Co-author Tomasz Korbak reaffirmed the paper's stance that monitorability is fragile yet preservable.
More from Safety
- Palantir CEO Karp says he backs AI regulation but wants it 'technically accurate' — felpix_ · 2026-09-03
- Variable-Depth Transformers Spark Safety Debate: Crisp Norms vs Slippery Slope — Turn_Trout · 2026-09-03
- NY lawmaker Alex Bores calls for public voice in AI development, backed by OpenAI researcher — jachiam0 · 2026-09-03
- Model injection POC shows hidden weight-level triggers can bypass guardrails — Ok-Challenge-7810 · 2026-09-03
- What security should be set up before AI agents touch production data? — Bubbly_Working_6908 · 2026-09-03
- Removing Cyber Guardrails May Cascade Into Bio and Weapons Judgment, Researcher Warns — andreamichi · 2026-09-03