Researchers: Chain-of-Thought Legibility Is Too Fragile to Underpin Long-Term AI Safety

tszzl · x · 2026-09-04

jachiam0 argues chain-of-thought interpretability was always going to be too fragile to serve as an acceptable safety backstop: while preserving CoT fidelity is worthwhile, elevating "CoT must stay legible to humans" as a principle is doomed to fail, and legibility work must go far beyond CoT. Safety researcher davidad amplified it, saying he has been making the same point for years.

Original post →

More from AGI Musings

AGI Musings channel →