Why CoT interpretability was never a real audit trail — and what a stronger safety substrate looks like

GaryMarcus · x · 2026-09-03

Gary Marcus quotes the view that 'chain of thought interpretability was always going to be so fragile as to be an unacceptable backstop for long-term AI safety,' but calls kicking away the scaffolding insane. @williamtp explains the architectural root: learned machinery lives in weights and much runtime computation hides in activations, so CoT's visible words are never a complete record of what determined the answer. The stronger destination: trust-critical reasoning in an inspectable symbolic substrate, where the trace is the reasoning — directly relevant to OpenAI's moves reducing CoT monitorability.

Related event: OpenAI's New Tech Reportedly Weakens CoT Monitorability, Sparking AI Safety Debate(32 posts)→

Original post →

More from Safety

Safety channel →