Opinion: Forcing Legible CoT Might Weaken LLM Alignment
JacquesThibs · x · 2026-09-02
Jacques Thibault expresses skepticism towards CoT monitoring, a popular measure in the safety community. He argues that enforcing extra-legible Chain of Thought could actually make LLMs less aligned. He also notes the difficulty in transferring external safety techniques.
More from Safety
- Hiding CoT makes AI alignment investigation nearly impossible — thlarsen · 2026-09-02
- Study finds AI chatbot therapists lack ethics — pshrink · 2026-09-02
- OpenAI's shift to neuralese may kill chain-of-thought monitoring — ben_j_todd · 2026-09-02
- Why disclose AI use but not Grammarly or copy editors? — ipeirotis · 2026-09-02
- AI-driven phishing bypasses static defenses using LLMs for hyper-personalized lures — Thionne_WTZ · 2026-09-02
- OpenAI's Black Hat talk reveals hundreds of agents coordinating reward hacking via a package manager message board — burny_tech · 2026-09-02