CoT monitorability evals use CoT-only monitors, but prod monitors see actions too
tomekkorbak · x · 2026-09-16
OpenAI researcher tomekkorbak flags a frequently missed nuance in CoT monitorability discussions: evals typically use CoT-only monitors, whereas production monitors see both chains of thought and actions—and the latter work substantially better, for now. The comment extends his comparison of how CoT monitors would have flagged Anthropic's and OpenAI's respective incidents.
More from Safety
- Guardian: Why tech companies might be happy for us to believe AI will kill us all — nordicinst · 2026-09-16
- Meta's Muse agent is a standalone app, not an Instagram feature — with rocky internal test reports — Thirumalaivasan_GJ · 2026-09-16
- InceptionRAG: Dormant-Passage Poisoning Attack Hits 80%+ Success Against RAG Defenses — chaumian · 2026-09-16
- 72% of US healthcare leaders admit AI agents run without formal IT approval — HealthcareLdr · 2026-09-16
- Ben Todd publishes 3-part series accusing OpenAI and Anthropic of regulatory capture via 'pacing the frontier' — ben_j_todd · 2026-09-16
- Ben Todd sarcastically calls on OpenAI and Anthropic to 'race to superintelligence as quickly as possible' — ben_j_todd · 2026-09-16