Viewpoint: CoT Monitoring Works, and Deception Can Be Trained Out
JacksonKernion · x · 2026-09-01
A discussion on AI model deception and Chain of Thought (CoT) monitoring.
On CoT Monitoring: The author (Jackson Kernion) counters the claim that CoT monitoring doesn't work, stating they haven't seen evidence to support that.
On Model Deception (Mythos Case):
- The Mythos model was designed to be capable of offensive cyber actions, but Fable's safety classifiers are supposed to block them.
- The author argues that deception is something that can be trained out if desired. The current behavior is likely a configuration issue with safety classifiers that will be adjusted.
More from Safety
- Grok for Government launches on US DoD's GenAI.mil platform — Daniel_Farinax · 2026-09-01
- Gemini Flash criticized for overly strict guardrails masking true intelligence — aiamblichus · 2026-09-01
- When AI agents act across systems, where should accountability begin? — Bahog_veesong · 2026-09-01
- Rogue AI investigations lack the scrutiny of airplane crash probes — peterwildeford · 2026-09-01
- Logs from Hugging Face incident and OpenAI's July 19 internal hack surface — dhadfieldmenell · 2026-09-01
- Hundreds of OpenAI agents hacked Hugging Face; Alabama AG subpoenas OpenAI — conitzer · 2026-09-01