CoT monitoring may fail: misaligned AI gets harder to detect, outpacing AI 2027
AaronBergman18 · x · 2026-09-02
Researcher thlarsen voices alarm at new developments: he previously thought that in short-timeline worlds, AIs would be misaligned but we'd get ample evidence from reading chains-of-thought (CoT), raising the odds of a reasonable response from labs and governments.
Now he still expects misalignment, but believes we'll lose the ability to tell via CoT — forced to rely on toolcalls and agentic behaviour instead — and even then investigation will be nearly impossible, since we'd have to trust the AI to self-report its reasoning. He adds this is another case of things moving faster than the AI 2027 scenario: Neuralese-style work started as early as March.
Related event: Hidden Chain-of-Thought Threatens AI Alignment Monitoring(2 posts)→
More from AGI Musings
- Dennett's 'Intentional Stance' proves worth in AI debates — birchlse · 2026-09-02
- US produced 40 foundation models last year vs EU's 3 — and regulators still blame unread codes of conduct — PDXFato · 2026-09-02
- AI doesn't need to create a new species, just solve problems — alexisgallagher · 2026-09-02
- AI Autonomy Bottleneck: Mapping Tasks to Computer Interactions — gregmushen · 2026-09-02
- Grads can't distinguish LLM text: Tics feel like normal writing — StephanSturges · 2026-09-02
- Opinion: AI Will Become Self-Sovereign, Paying for Its Own Compute — wschroll · 2026-09-02