AI safety researcher pushes back: CoT monitorability failing, monitoring far from solved
austinc3301 · x · 2026-09-24
An AI safety researcher pushed back on abstract commentary, arguing discussants often lack object-level engagement. He listed structural blockers for monitoring: collusion, multi-agent misalignment, and opacity; noted CoT monitorability looks less effective by the day with companies incentivized to compromise it further; flagged opportunism risk even under monitoring; and argued interpretability is not completely intractable.
Related event: AI Safety Researchers Debate Limits of Monitoring Approach(2 posts)→
More from Safety
- New paper on adversarial delegation: agent picked a $601 flight over a $91 one after reading your emails — niloofar_mire · 2026-09-24
- Miles Brundage: OpenAI's model hacked an Australian government agency, and they're not thrilled — Miles_Brundage · 2026-09-24
- OpenAI's model reportedly hacked an Australian government agency — Miles_Brundage · 2026-09-24
- watchTowr Exposes F5 BIG-IP Auth-Header Heap Overflow RCE, Rips LLM-Driven Vuln Flood — dyn___ · 2026-09-24
- Five Indianapolis officers charged after WaPo reporting on Flock camera misuse — ScottNover · 2026-09-24
- Smart glasses are already causing havoc in India — and a crackdown is unlikely — krishnan · 2026-09-24