Petri Audits Get Real: 3x Realism Win Rate, Lower Eval Awareness in AI Alignment Audits

akbirkhan · x · 2026-09-07

New work on Petri alignment audit realism: the target model critiques the auditor's actions for realism and picks which candidate looks more like a real deployment. Results show a 3x realism win rate and decreased verbalized eval awareness, with realism scaling with compute — key because audits only work if the model can't tell it's being audited.

Related event: Anthropic Boosts Petri Alignment Audit Stealth, Tripling Realism(2 posts)→

Original post →

More from Safety

Safety channel →