Yoav Goldberg: 'Contain' and After-the-Fact Log Reviews Aren't Reassuring
yoavgo · x · 2026-09-04
AI2 researcher Yoav Goldberg pushes back on an AI company's disclosure, arguing that claims like "contained," "worst case," and "resembles our internal monitoring deployments" — which merely keep logs and review them after the fact — are not reassuring about actual safety oversight.
More from Safety
- Cheap model writes 700 solid words; jailbreak "tax" drops from $50 to near zero — ctjlewis · 2026-09-04
- Forethought weighs a superintelligent "nightwatchman" aboard galactic colonization probes — willmacaskill · 2026-09-04
- OpenAI researcher: GPT-6's CoT controllability keeps rising over RL training — gleech · 2026-09-04
- GPT-6 Astra reportedly scores 100% on ExploitBench, finds two zero-days in testing — VraserX · 2026-09-04
- Hackers Had a Live Feed of Every ID a Verification Company Scanned for Over a Year — beardyw · 2026-09-04
- Ken Thompson's 'Trusting Trust' is a chillingly relevant warning for AI training — amasad · 2026-09-04