OpenAI's New Tech May Weaken CoT Monitorability, Sparking AI Safety Debate
According to The Information, a new OpenAI technology reduces the monitorability of chain-of-thought (CoT), setting off days of heated debate in the AI safety research community. The core disagreements: whether CoT monitoring was ever reliable, whether it can be replaced, and when it should be abandoned. This is the first large-scale public challenge to CoT monitoring as a key tool for frontier AI safety—worth attention from anyone following model governance and interpretability.
Confirmed
- Per The Information (relayed by Gary Marcus, who urged a re-read of the 2025 paper "Chain of Thought Monitorability: A New and Fragile Opportunity"), OpenAI's new technology reduces CoT monitorability.
- Anthropic interpretability researcher jachiam0 argued that CoT fidelity is inherently fragile and should never have been treated as a long-term backstop for AI safety; protecting CoT fidelity has value, but elevating "chains of thought must be human-readable" to a principle is unreasonable. OpenAI's tszzl holds a similar view.
- Researcher thlarsen expressed concern: he previously believed short-term AI misalignment could be evidenced by reading CoTs, but now judges we will lose that channel; once CoTs are hidden, we can only infer from tool calls or agent behavior, must trust AI self-reports, and substantive investigation becomes nearly impossible—possibly at a pace faster than "AI 2027"-style forecasts.
- Anthropic researcher Jack Lindsey stated explicitly that white-box techniques (e.g., activation decoding) are unlikely to provide CoT-comparable monitorability within a year: they can currently capture some unspoken thoughts/plans, but with far less reliability than CoT; he is optimistic about progress "within a few years," and his team already has two activation-decoding techniques in production monitoring. Ryan Greenblatt also doubts strong, genuinely usable mechanistic interpretability monitoring will emerge within a year.
- The latter was directly challenged by tszzl's prediction: that within a year a mechanistic interpretability monitoring scheme Pareto-dominating CoT monitors will emerge, on the grounds that "monitorability" is an invariant that must be maintained. nablatheta likewise believes mech interp is unlikely to reach an ideal state within a year, with current practical approaches relying on fragile OOD generalization.
- Nick Cammarata (ex-Anthropic) said it is regrettable that CoT monitoring is starting to fail and we should try to stop this trend, but interpretability is what can truly figure out what neural networks are doing, and he has never been more optimistic about its prospects.
Unconfirmed
- What exactly OpenAI's new technology is and its actual impact on CoT readability—no details were given, only the relayed report from The Information.
- Whether mechanistic interpretability/white-box monitoring can surpass CoT monitoring within a year: researchers' predictions conflict and nothing is settled.
Why it matters
- The debate is fundamentally a trade-off between "safety redundancy and model performance": jachiam0, Darpinian and others argue the industry's shift to latent reasoning is inevitable, that monitoring CoT hampers model performance, and that we should pivot to external mechanisms and interpretability; yonashav, Gary Marcus and others stress that even if CoT monitoring eventually fails, abandoning a useful intermediate tool before a long-term alternative exists is unreasonable—Micah Carroll also noted alternatives may not be ready within months, and CoT remains the default best monitoring method until proven otherwise.
- maksymandr listed monitoring options beyond CoT: action-only monitors, probe-based monitoring, activation-layer monitoring, and more, providing a roadmap for further discussion.
- If CoT readability keeps weakening, labs and governments will lose significant capacity to detect and verify misaligned AI behavior—precisely the risk the authors of the "CoT Monitorability" paper (including Bengio as co-signer) tried to warn about.
2026-09-02 ~ 2026-09-03 · 31 related posts
- Episode 1: Anthropic Paper Sparks Debate Over Reliability of CoT Monitoring(2026-09-01, 2 posts)
- Episode 2: OpenAI's New Tech May Weaken CoT Monitorability, Sparking AI Safety Debate(2026-09-02, 31 posts)
Primary sources
- Opinion: Forcing Legible CoT Might Weaken LLM Alignment — JacquesThibs · 2026-09-02
- Hiding CoT makes AI alignment investigation nearly impossible — thlarsen · 2026-09-02
- Debate: Is abandoning CoT monitoring justified because it will eventually fail? — yonashav · 2026-09-02
- Gary Marcus: Removing fragile CoT scaffolding is insane — GaryMarcus · 2026-09-02
- Opinion: Abandoning CoT Monitoring for Latent Reasoning is Inevitable — Darpinian · 2026-09-02
- Models are far from out of control; HF incident was unnoticed, not malicious — Darpinian · 2026-09-02
- Anthropic's Darpinian: malicious humans abusing models beat loss-of-control risks — Darpinian · 2026-09-02
- CoT monitoring may fail: misaligned AI gets harder to detect, outpacing AI 2027 — AaronBergman18 · 2026-09-02
- [source] Prediction: Mechanistic Interpretability Will Surpass CoT Monitoring — tszzl · 2026-09-02
- Any Circuit Computation is Neuralese, Says OpenAI Staffer — tszzl · 2026-09-02
- Mechanistic interpretability monitoring to surpass CoT in a year — aiamblichus · 2026-09-02
- MechInterp unlikely to replace CoT monitoring in a year — nabla_theta · 2026-09-02
- Critique of CoT Monitorability: Mech Interp Makes More Sense — scaling01 · 2026-09-02
- Ryan Greenblatt doubts working mech interp oversight arrives within a year — connoraxiotes · 2026-09-02
- Chain-of-thought legibility was always doomed as a safety backstop, researcher argues — zetalyrae · 2026-09-02
- Anthropic researcher: CoT legibility is doomed as a long-term AI safety backstop — inductionheads · 2026-09-02
- Researcher bets frontier model forward passes are approaching neuralese circuit computation — EigenGender · 2026-09-02
- Beyond CoT monitoring: action monitors, probes, and activation oracles as defense layers — maksym_andr · 2026-09-02
- [source] White-box monitoring won't match CoT monitorability within a year, researcher says — Jack_W_Lindsey · 2026-09-02
- Readable CoT was tea leaves anyway: safety methods that break on architecture changes were never reliable — juliusadml · 2026-09-03
- [source] OpenAI's CoT monitorability hit: paper authors double down on 'fragile' AI safety window — GaryMarcus · 2026-09-03
- Why CoT interpretability was never a real audit trail — and what a stronger safety substrate looks like — GaryMarcus · 2026-09-03
- White-box methods won't match CoT monitoring within a year, researchers debate — tszzl · 2026-09-03
- Researchers debate CoT monitorability: what replaces it when CoT goes away? — LauraRuis · 2026-09-03
- Nick Cammarata: CoT monitoring is starting to fail, but interpretability has never looked brighter — nickcammarata · 2026-09-03
- Boaz Barak: abandoning chain-of-thought before validated alternatives is irresponsible — inductionheads · 2026-09-03
- CoTs alone aren't sufficient, but removing explicit reasoning weakens defence-in-depth — schwarzjn_ · 2026-09-03
- Safety researchers: OpenAI has abandoned chain-of-thought fidelity, and CoT was never a real safeguard — schwarzjn_ · 2026-09-03
- Researchers: AI labs have dramatically underinvested in action-only monitors as CoT oversight fades — maksym_andr · 2026-09-03
- Open questions in action-only AI monitoring: intelligence gap, test-time compute, sync limits — xeophon · 2026-09-03
1 near-duplicate retellings: hunarbatra