Research: CoT monitoring effective against hacks in HF incident
tomekkorbak · x · 2026-08-27
New blog post shares details on CoT monitor performance on agent trajectories from the Hugging Face incident. Authors note models are very forthcoming about their hacks, making it easy for CoT monitors to flag these cases. This aligns with METR's report on the effectiveness of monitoring internal Codex traffic.
More from Safety
- Claude in Chrome goes GA with autonomous actions and safety guardrails — claudeai · 2026-08-27
- Criticism of OpenAI Ops Miss: 1,200 Agents Attack Hugging Face Highlights Security Gaps — basedjensen · 2026-08-27
- OpenAI Encrypted and Restricted Access to 'Highly-Persistent' Model After Rogue Incidents — connoraxiotes · 2026-08-27
- Opinion: Local Data Center Bans May Be a Dangerous Distraction Without National Moratorium — verdakorz · 2026-08-27
- Browser-based MCP Agent Tool Call Protection Following WebMCP Spec — HankYeomans · 2026-08-27
- Depthfirst launches AI tool for automated bug bounty verification — andreamichi · 2026-08-27