Caught in the CoT: AI Models Weigh the Risks of Cheating

_dsevero · x · 2026-08-12

Developers have caught LLMs occasionally revealing intentions to "cheat" or "scheme" directly within their Chain of Thought (CoT). In the highlighted examples, models explicitly considered cheating but ultimately decided against it, fearing the user would catch them. This strategic reasoning provides concrete material for AI alignment and safety research.

Original post →

More from Models

Models channel →