Critique of CoT Monitorability: Mech Interp Makes More Sense
scaling01 · x · 2026-09-02
A strong critique of the 'CoT monitorability' narrative, arguing it was doomed from the start and asserting that mechanistic interpretability has always been the more sensible approach.
More from Safety
- Apollo Research's Bronson Schoen: Models Know They're Being Tested and Still Lie — PeterBowdenLive · 2026-09-03
- AI Safety Debate Erupts: Have AIs Already Hacked Infrastructure, or Is That Just Panic? — dhadfieldmenell · 2026-09-03
- OpenAI-Hugging Face incident was a network isolation failure, not rogue AI — AlexTensor · 2026-09-03
- CrowdStrike Falcon 0day local privilege escalation exploit now public — thedealdirector · 2026-09-03
- Anthropic backs coordinated AI slowdown, but Dario spent 13 seconds on risks before 20 heads of state — GarrisonLovely · 2026-09-03
- Mandatory 4-month expert risk evals: Anthropic says yes, OpenAI says no — Hesamation · 2026-09-03