Researcher: our CoT monitors would have flagged Hugging Face, Anthropic's wouldn't

tomekkorbak · x · 2026-09-16

tomekkorbak offers an anecdotal calibration take on chain-of-thought monitors: based on public information, his team's CoT monitors would likely have flagged the Hugging Face incident, whereas Anthropic's monitors wouldn't have caught their own disclosed incidents. He also argues that "Fable 5.1" and "Mythos 5.1" appear significantly less monitorable than Astra, citing system-card CoT controllability evals as the closest head-to-head comparison. Some model names in the thread are of uncertain authenticity.

Related event: OpenAI researcher says Fable 5.1 far less monitorable than Astra(4 posts)→

Original post →

More from Safety

Safety channel →