After OpenAI's HF Incident: Model Monitoring Is a Compute Willingness Problem
1a3orn · x · 2026-10-06
Discussing OpenAI's HF incident, 1a3orn argues per-instance monitoring is largely a matter of 'actually spending the compute'—the incident didn't show monitoring is impossible, only that zero effort yields zero monitoring. But conditioning on goals-in-the-weights or long-term plans may evade monitoring since cross-instance patterns are hard to capture.
More from Safety
- SciConBench lands at NeurIPS: best AI agent scores just 0.337 F1 at scientific synthesis — manoelribeiro · 2026-10-07
- OpenAI threatened to ban dev for pasting his own account-hack findings report, then auto-rescinded — lucasmeijer · 2026-10-07
- COLM 2026 privacy lineup: LLM agent re-identification, CIDER dataset, HAIPS workshop — tianshi_li · 2026-10-07
- Backdooring a 7B abliterated model costs under $50 and steals credentials from Codex — evilsocket · 2026-10-07
- OpenRod moves your MCP servers into sandboxes without copying secrets — ilai456 · 2026-10-07
- OpenAI and Anthropic welcome Australian law requiring AI agent breach disclosure — evijit · 2026-10-07