Researcher: our CoT monitors would have flagged Hugging Face, Anthropic's wouldn't
tomekkorbak · x · 2026-09-16
tomekkorbak offers an anecdotal calibration take on chain-of-thought monitors: based on public information, his team's CoT monitors would likely have flagged the Hugging Face incident, whereas Anthropic's monitors wouldn't have caught their own disclosed incidents. He also argues that "Fable 5.1" and "Mythos 5.1" appear significantly less monitorable than Astra, citing system-card CoT controllability evals as the closest head-to-head comparison. Some model names in the thread are of uncertain authenticity.
Related event: OpenAI researcher says Fable 5.1 far less monitorable than Astra(4 posts)→
More from Safety
- Reported FTC probe: OpenAI may face liability over Hugging Face incident — AIFlow_ML · 2026-09-16
- Bill Kristol: AI guardrails without enforcement, liability and penalties aren't real guardrails — Miles_Brundage · 2026-09-16
- Anthropic and OpenAI spend only ~1% of capex on AI safety, analysis shows — AlexTensor · 2026-09-16
- 53 MCP servers scanned: 36% graded D/F, mostly for over-permissioned scope — BrilliantSecret143 · 2026-09-16
- CMU's Decoy Direction Optimization blocks refusal-ablation attacks at 30-450x lower cost — CarnegieMellonU · 2026-09-16
- US voters oppose AI data centers 57%-71%, while DOJ backs OpenAI's fair-use defense vs NYT — emmanuelvivier · 2026-09-16