Anthropic researcher: judge AI labs by safety outcomes, not stated policies

kipperrii · x · 2026-08-31

Anthropic researcher Ethan Perez argues that AI labs should be evaluated on outcomes, as this sets good incentives. The factors that matter most in preventing major safety incidents are largely non-public: who gets hired and promoted, org culture, vetting of training environments and data, how cautiously researchers explore risky ideas, cross-talk between safety/security/capabilities teams, and compensation incentives. These are the least legible to outside observers, making careful examination of incident track records essential. He adds that Anthropic's and OpenAI's incidents differ in both kind and severity (single-rollout incidents vs. multi-week problems).

Original post →

More from Safety

Safety channel →