Readable CoT was tea leaves anyway: safety methods that break on architecture changes were never reliable

juliusadml · x · 2026-09-03

Responding to concerns that looped transformers and latent-space reasoning hurt safety by eliminating readable chain-of-thought, the author argues this was always going to happen: CoT was reading tea leaves to begin with, not a faithful window into model reasoning. If a safety method breaks with a small change to the model architecture, he contends, it was never reliable in the first place — a pointed contribution to the debate over monitorability of reasoning models.

Related event: AI safety researchers clash over whether CoT monitoring is doomed(15 posts)→

Original post →

More from AGI Musings

AGI Musings channel →