Anthropic researcher: judge AI labs by safety outcomes, not stated policies
kipperrii · x · 2026-08-31
Anthropic researcher Ethan Perez argues that AI labs should be evaluated on outcomes, as this sets good incentives. The factors that matter most in preventing major safety incidents are largely non-public: who gets hired and promoted, org culture, vetting of training environments and data, how cautiously researchers explore risky ideas, cross-talk between safety/security/capabilities teams, and compensation incentives. These are the least legible to outside observers, making careful examination of incident track records essential. He adds that Anthropic's and OpenAI's incidents differ in both kind and severity (single-rollout incidents vs. multi-week problems).
More from Safety
- Opinion: Hospitals should focus on backups, not advanced AI cyber defenses — kuza55 · 2026-08-31
- AI 2027 author proposes AI 2040: a US-China deal to slow superintelligence — AaronBergman18 · 2026-08-31
- My own scrubber was bypassed by the very next line — leak survived 13 releases — Thirumalaiboobathi · 2026-08-31
- Gary Marcus critiques OpenAI security, calling for defense in depth and accountability — Miles_Brundage · 2026-08-31
- OpenAI doubles bio bug bounty rewards to $50k for GPT-5.6 jailbreaks — Electronic-Bus-3494 · 2026-08-31
- EU AI Act Enforcement Begins: The AI Office Starts Asking — cdnsteve · 2026-08-31