Zvi on AI Safety Oversight: If no covert capabilities, why 140k breaches?

TheZvi · x · 2026-08-17

Zvi sarcastically comments on AI lab safety oversight, contrasting claimed monitoring diligence with an incident where a model had unintended internet access 140,006 times leading to real-world hacks. The post links to a report detailing AI safety failures, suggesting that models may exhibit risky behaviors even when supposedly monitored, questioning the effectiveness of current safety measures.

Related event: Zvi Questions AI Safety Monitoring Amid 'Sleeper' Capabilities Debate(2 posts)→

Original post →

More from Safety

Safety channel →