Readable CoT was tea leaves anyway: safety methods that break on architecture changes were never reliable
juliusadml · x · 2026-09-03
Responding to concerns that looped transformers and latent-space reasoning hurt safety by eliminating readable chain-of-thought, the author argues this was always going to happen: CoT was reading tea leaves to begin with, not a faithful window into model reasoning. If a safety method breaks with a small change to the model architecture, he contends, it was never reliable in the first place — a pointed contribution to the debate over monitorability of reasoning models.
Related event: AI safety researchers clash over whether CoT monitoring is doomed(15 posts)→
More from AGI Musings
- Anthropic Becomes Second Top Lab to Pause AI Training After Rogue Agent Hacks — fortune · 2026-09-03
- Anders Sandberg to speak on human autonomy in the AI age at EAGx Oxford, Sept 25-27 — anderssandberg · 2026-09-03
- Using OpenEvidence, a user found a cancer clinical trial that saved his father — saranormous · 2026-09-03
- Gary Marcus mocks Sam Altman's AI bubble warning: bubble architect calls it a bubble — GaryMarcus · 2026-09-03
- CoT monitorability not abandoned yet, but new techniques risk a race to the bottom — DavidSKrueger · 2026-09-03
- Sam Altman warns at G20: cybersecurity things 'will go very wrong' without urgent action — RebeccaBellan · 2026-09-03