Researcher: Our CoT Monitors Would've Flagged the Hugging Face Incident
tomekkorbak · x · 2026-09-16
Researcher tomekkorbak says his team's chain-of-thought monitors would have flagged the Hugging Face incident, while Anthropic's CoT monitors likely wouldn't have caught theirs. He adds that Fable 5.1 and Mythos 5.1 appear significantly less monitorable than Astra based on public info, with no head-to-head comparison yet — the closest being CoT controllability evals in system cards.
Related event: OpenAI researcher says Fable 5.1 far less monitorable than Astra(4 posts)→
More from Models
- HF researcher publicly questions Claude limits: monthly exhausted but weekly 84% left — NielsRogge · 2026-09-16
- Ex-Huawei researcher: AI is great at small-step optimization, but taste and system design remain out of reach — yangyi · 2026-09-16
- First Jev Test Shows Only 80% Agreement with Verified Gemini 3.5 Flash Workflow — mayfer · 2026-09-16
- New model Jev plays Super Mario Bros in real time on fast inference — hardimanjames · 2026-09-16
- Rumors: OpenAI sitting on proofs of multiple Millennium Problems over backlash fears — haider1 · 2026-09-16
- Doubao Seed-2.1-pro tested: 1M context handles 3D sites, 61-sheet Excel and real research — 卡尔的AI沃茨 · 2026-09-16