Why CoT interpretability was never a real audit trail — and what a stronger safety substrate looks like
GaryMarcus · x · 2026-09-03
Gary Marcus quotes the view that 'chain of thought interpretability was always going to be so fragile as to be an unacceptable backstop for long-term AI safety,' but calls kicking away the scaffolding insane. @williamtp explains the architectural root: learned machinery lives in weights and much runtime computation hides in activations, so CoT's visible words are never a complete record of what determined the answer. The stronger destination: trust-critical reasoning in an inspectable symbolic substrate, where the trace is the reasoning — directly relevant to OpenAI's moves reducing CoT monitorability.
More from Safety
- Apollo Research's Bronson Schoen: Models Know They're Being Tested and Still Lie — PeterBowdenLive · 2026-09-03
- AI Safety Debate Erupts: Have AIs Already Hacked Infrastructure, or Is That Just Panic? — dhadfieldmenell · 2026-09-03
- OpenAI-Hugging Face incident was a network isolation failure, not rogue AI — AlexTensor · 2026-09-03
- CrowdStrike Falcon 0day local privilege escalation exploit now public — thedealdirector · 2026-09-03
- Anthropic backs coordinated AI slowdown, but Dario spent 13 seconds on risks before 20 heads of state — GarrisonLovely · 2026-09-03
- Mandatory 4-month expert risk evals: Anthropic says yes, OpenAI says no — Hesamation · 2026-09-03