OpenAI Reportedly Testing Tech That Obscures Chain-of-Thought, Sparking AI Safety Debate
The Information reported that OpenAI is experimenting with a new technique that shows less of the model's "Chain of Thought" in its outputs, making it harder to monitor. Gary Marcus immediately published a warning that the move crosses an AI safety red line, triggering a fierce debate in the AI safety community—ranging from doomsday readings like "chain of thought is dead" to sober analyses of the underlying technology.
Confirmed
- The Information reported that OpenAI is exploring new technology to reduce the exposure of the model's thought process (as relayed by m4, m8).
- Gary Marcus publicly and strongly criticized the plan, arguing that sacrificing CoT monitorability for possibly tiny performance gains is "begging for disaster," and noted that former OpenAI safety researcher Steven Adler also expressed concerns (m1, m4, m8).
- Geoffrey Irving pushed back on the argument that "using recurrent structures is still safe," contending that shallow circuits cannot guarantee safety (m6).
- Former OpenAI employee Joshua defended hiding CoT outputs, saying chain of thought is too fragile and doomed to fail as a long-term safety fallback and that "CoT must be human-readable" should not be enshrined as a principle; GarrisonLovely criticized his flip-flopping stance and attacked the reporting outlet (m5).
Unconfirmed
- Whether OpenAI has actually adopted a Looped Transformer / Recurrent Depth architecture remains based on a single media report, with no official confirmation from OpenAI. maxpaperclips explicitly noted that the rumors are hearsay (m2, m7).
- Whether the technology truly makes models "harder to monitor" is disputed: maxpaperclips argues recurrent depth is essentially equivalent to making the network 10x deeper with early exit—not some novel unmonitorable black magic, just compute overhead in disguise (m2, m7). sjgadler also believes the "neural language, chain of thought is dead" narrative may be an overinterpretation (m3).
Why it matters
- CoT monitorability, though limited, is seen as a key clue for preventing severe outcomes; the safety camp worries that if OpenAI abandons this clue first, it could trigger a race to the bottom—sjgadler fears other companies will follow suit and drop safety constraints, believing OpenAI broke the pact first, and "someone else breaking it" is no excuse for your own betrayal (m1, m3).
- At its core, the debate reflects a split over whether "CoT readability as a safety principle" should be the path: one side treats it as a non-negotiable line of defense, while the other (e.g., Joshua) sees it as fragile and unreliable, not worth being a long-term safety mechanism (m5, m6).
2026-09-02 ~ 2026-09-02 · 10 related posts
Primary sources
- Gary Marcus: OpenAI's new technique could destroy chain-of-thought monitorability — GaryMarcus ·
- Debunking Looped Transformer Hype: It's Just More FLOPs, Not AGI — max_paperclips ·
- Geoffrey Irving on looped transformers: "Fake bounds" don't ensure safety — geoffreyirving ·
- [source] Debunking Looped Transformer Hype: It's Just More FLOPs, Not AGI — max_paperclips · 2026-09-02
- Ex-OpenAI Staffer Defends 'Hidden Thought' Strategy Amid Criticism — GarrisonLovely · 2026-09-02
- Gary Marcus Warns OpenAI's Plan to Hide Reasoning is a Safety Redline — GaryMarcus · 2026-09-02
- Experts warn OpenAI's opaque CoT moves could trigger dangerous safety race — sjgadler · 2026-09-02
- Debating OpenAI's "Recurrent Depth": Not harder to monitor, just deeper — max_paperclips · 2026-09-02
- [source] Geoffrey Irving on looped transformers: "Fake bounds" don't ensure safety — geoffreyirving · 2026-09-02
- GaryMarcus warns OpenAI reportedly sacrificing CoT monitorability for performance — AndyMasley · 2026-09-02
- Recurrent Activations Raise AI Monitoring Challenges — RyanGreenblatt · 2026-09-02
2 near-duplicate retellings: GaryMarcus · GaryMarcus