CoT Monitoring Fragile and Unreliable; Action-Only Monitors Deserve Attention
maksym_andr · x · 2026-08-16
maksymandr argues that CoT monitoring, called a fragile opportunity for AI safety, already looks fragile and unreliable. He suggests preparing for a future without relying on CoT monitoring, and that action-only monitors deserve a closer look.
In ResearchArena, CoT monitors can be misled if they rely too much on CoT, as shown in an example where an action-only monitor is correct while a CoT monitor fails.
More from Safety
- Deepfake Australian PM used in celebrity scams causing $7.4m in losses — nordicinst · 2026-08-16
- AI chatbot's pesticide advice wipes out 25 acres of Chinese farmer's sesame seedlings — luisdans · 2026-08-16
- CFC Rule Tested: Can Simple Control Stop LLMs from Hallucinating Decisions? — Plastic-Cell-4497 · 2026-08-16
- Critics Argue Anthropic's Watermarking Scheme Fuels Global Surveillance — nptacek · 2026-08-16
- User drops Anthropic over watermarking, arguing tech always has an exit from surveillance — StewartalsopIII · 2026-08-16
- Naval suggests legislation: Open models required if trained on open web — rohanpaul_ai · 2026-08-16