New paper makes Petri alignment audits stealthier: 3x realism win rate, less eval awareness

EthanJPerez · x · 2026-09-07

A new paper shared by Anthropic's Ethan Perez argues alignment audits only work if the model can't tell it's being audited. The team made Petri audits far more realistic, tripling the realism win rate and reducing verbalized eval awareness, making audit results more trustworthy.

Related event: Anthropic Boosts Petri Alignment Audit Stealth, Tripling Realism(2 posts)→

Original post →

More from Safety

Safety channel →