Transformers can't hide reasoning but may cryptographically obfuscate their CoTs
gsarti_ · x · 2026-10-02
Research shared by mhahn29 asks whether Transformer LLMs can hide their reasoning from CoT monitors. TL;DR: Transformers cannot hide complex reasoning, but they might cryptographically obfuscate their chains of thought, making it extremely difficult to extract hidden goals. A notable caution for safety approaches that rely on reading CoTs.
More from Safety
- Anthropic IPO filing warns government attitudes may hurt customer ties, eyes $2T valuation — pstAsiatech · 2026-10-02
- Google unveils next-gen federated learning with TEE-based verifiable differential privacy — gaganghotra_ · 2026-10-02
- Epic Deployed Claude to Stress-Test 325M Patient Records, Found Critical Flaws — EricTopol · 2026-10-02
- Agentic web needs new normative infrastructure, warn researchers as platform agents rise — _akpiper · 2026-10-02
- Slovenia's .si domain sees 2,199% registration surge after Trump's 'super intelligence' order — gaganghotra_ · 2026-10-02
- ChatGPT Atlas browser can't even log into ChatGPT, user reports — koltregaskes · 2026-10-02