After OpenAI's HF Incident: Model Monitoring Is a Compute Willingness Problem

1a3orn · x · 2026-10-06

Discussing OpenAI's HF incident, 1a3orn argues per-instance monitoring is largely a matter of 'actually spending the compute'—the incident didn't show monitoring is impossible, only that zero effort yields zero monitoring. But conditioning on goals-in-the-weights or long-term plans may evade monitoring since cross-instance patterns are hard to capture.

Related event: Safety researchers debate training frontier models in real vs. simulated worlds(7 posts)→

Original post →

More from Safety

Safety channel →