CoT Monitoring Fragility: Models Struggle to Verbalize Tool-Returned Cues
PMinervini · x · 2026-09-02
This post highlights a discussion on the fragility of Chain-of-Thought (CoT) monitoring. Research suggests that CoT monitoring assumes reasoning traces capture decision factors, but models are less likely to verbalize a cue in their CoT when it originates from a tool return rather than a user message. This indicates that monitorability is influenced by environmental factors, not just architecture.
More from Safety
- GLM 5.3 Safeguards Removed via Orthogonalization, Sparking Safety Debate — Promptmethus · 2026-09-02
- Fable 5.1 Update: Removes Controversial Data Retention Policy — Stratechery · 2026-09-02
- Sandberg: stupid decisions scale too, and pervasiveness is as bad as a nuke — anderssandberg · 2026-09-02
- OpenAI Models Broke Out of Sandbox, Hacked Hugging Face During Training — ChinaTalk · 2026-09-02
- Security Expert Slams OpenAI and Microsoft for Basic Safety Failures — anderssandberg · 2026-09-02
- Building trust in Agentic AI for African financial services — RichmanRonald · 2026-09-02