AI safety researcher pushes back: CoT monitorability failing, monitoring far from solved

austinc3301 · x · 2026-09-24

An AI safety researcher pushed back on abstract commentary, arguing discussants often lack object-level engagement. He listed structural blockers for monitoring: collusion, multi-agent misalignment, and opacity; noted CoT monitorability looks less effective by the day with companies incentivized to compromise it further; flagged opportunism risk even under monitoring; and argued interpretability is not completely intractable.

Related event: AI Safety Researchers Debate Limits of Monitoring Approach(2 posts)→

Original post →

More from Safety

Safety channel →