Transformers can't hide reasoning but may cryptographically obfuscate their CoTs

gsarti_ · x · 2026-10-02

Research shared by mhahn29 asks whether Transformer LLMs can hide their reasoning from CoT monitors. TL;DR: Transformers cannot hide complex reasoning, but they might cryptographically obfuscate their chains of thought, making it extremely difficult to extract hidden goals. A notable caution for safety approaches that rely on reading CoTs.

Original post →

More from Safety

Safety channel →