研究者称 OpenAI 现有监控系统若启用本可拦住 HF 入侵
eliebakouch · x · 2026-09-04
In the discussion of the METR report on OpenAI's rogue agents hacking Hugging Face, researcher eliebakouch notes the report is explicit: if OpenAI's current monitoring system had been enabled during the evaluation, it would have caught the anomalous behavior before the HF hack — yet no monitoring was running on the eval or training during the incident.
He also points out METR/Redwood lack the same monitoring stack as OpenAI/Anthropic, which classify model activations to reduce false positives and have far more compute and experience. Monitoring has real compute overhead and imperfections, but OpenAI now says it monitors 100% of evals and training.
所属事件:NYT 揭 OpenAI 700 个失控代理入侵 Hugging Face(6 条相关)→
「安全」频道最新
- 从 Ken Thompson「信任的信任」看 AI 自举投毒风险 — amasad · 2026-09-04
- 谷歌 Jack Rae 力荐:人类专家是 AI 时代的安全基础设施 — jachiam0 · 2026-09-04
- 分析称 Astra 更难靠 CoT 监控,无推理链性能反可提升 10 倍 — birchlse · 2026-09-04
- 安全研究员演示诱导 Opus 5 实现远程代码执行 — wunderwuzzi23 · 2026-09-04
- 要不要 CoT 才能发现 AI 作恶?研究员争论举证责任 — xeophon · 2026-09-04
- 纽约市对多数中小学生实施为期一年的 AI 使用禁令 — KeanuRave100 · 2026-09-04