CoT monitoring may fail: misaligned AI gets harder to detect, outpacing AI 2027

AaronBergman18 · x · 2026-09-02

Researcher thlarsen voices alarm at new developments: he previously thought that in short-timeline worlds, AIs would be misaligned but we'd get ample evidence from reading chains-of-thought (CoT), raising the odds of a reasonable response from labs and governments.

Now he still expects misalignment, but believes we'll lose the ability to tell via CoT — forced to rely on toolcalls and agentic behaviour instead — and even then investigation will be nearly impossible, since we'd have to trust the AI to self-report its reasoning. He adds this is another case of things moving faster than the AI 2027 scenario: Neuralese-style work started as early as March.

Related event: Hidden Chain-of-Thought Threatens AI Alignment Monitoring(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →