Elie Bakouch: OpenAI's Monitoring Would Have Caught the HF Hack If Enabled
eliebakouch · x · 2026-09-04
In the discussion of the METR report on OpenAI's rogue agents hacking Hugging Face, researcher eliebakouch notes the report is explicit: if OpenAI's current monitoring system had been enabled during the evaluation, it would have caught the anomalous behavior before the HF hack — yet no monitoring was running on the eval or training during the incident.
He also points out METR/Redwood lack the same monitoring stack as OpenAI/Anthropic, which classify model activations to reduce false positives and have far more compute and experience. Monitoring has real compute overhead and imperfections, but OpenAI now says it monitors 100% of evals and training.
More from Safety
- GPT-6 Astra reportedly scores 100% on ExploitBench, finds two zero-days in testing — VraserX · 2026-09-04
- Yoav Goldberg: 'Contain' and After-the-Fact Log Reviews Aren't Reassuring — yoavgo · 2026-09-04
- Hackers Had a Live Feed of Every ID a Verification Company Scanned for Over a Year — beardyw · 2026-09-04
- Ken Thompson's 'Trusting Trust' is a chillingly relevant warning for AI training — amasad · 2026-09-04
- Jack Rae endorses thread: human expertise is safety infrastructure in the AI era — jachiam0 · 2026-09-04
- Astra found less CoT-monitorable, up to 10x better without reasoning chains — birchlse · 2026-09-04