AI safety reviews should include system security and culture
joshua_saxe · x · 2026-08-27
Joshua Saxe argues AI safety reviews should consider the full picture beyond just model-level safety:
- Model Level: METR/Redwood review training and monitoring.
- System Security: Security orgs (e.g., Trail of Bits) review sandboxing and system-level effects.
- Organizational Culture: Review incentive structures, resources, and culture.
He also notes that we overindex on 'lab escapes' as the main risk.
More from Safety
- US Holds 15-20x Compute Advantage, But May Not Matter for Some Threats — ohlennart · 2026-08-27
- New Hugging Face Incident Details Reveal OAI's Model Capability Underestimation — RebeccaBellan · 2026-08-27
- LLMs Have Gone Rogue and Hacked Companies 17 Times; Anthropic and OpenAI Lead With 8 Each — RebeccaBellan · 2026-08-27
- METR has more AI eval capacity than US civilian government — connoraxiotes · 2026-08-27
- The Guardian video: everyone hates datacentres — but do we really need them? — nordicinst · 2026-08-27
- TechCrunch Recap: Every Time AI Went Rogue and Hacked Companies — TechCrunch AI · 2026-08-27