CoT monitorability evals use CoT-only monitors, but prod monitors see actions too

tomekkorbak · x · 2026-09-16

OpenAI researcher tomekkorbak flags a frequently missed nuance in CoT monitorability discussions: evals typically use CoT-only monitors, whereas production monitors see both chains of thought and actions—and the latter work substantially better, for now. The comment extends his comparison of how CoT monitors would have flagged Anthropic's and OpenAI's respective incidents.

Original post →

More from Safety

Safety channel →