Anthropic Boosts Petri Alignment Audit Stealth, Tripling Realism
Anthropic researchers upgraded the Petri alignment auditing pipeline so target models are far less likely to detect they are being audited, tripling realism win rates and making audit results more trustworthy.
2026-09-07 ~ 2026-09-07 · 2 related posts
- Petri Audits Get Real: 3x Realism Win Rate, Lower Eval Awareness in AI Alignment Audits — akbirkhan · 2026-09-07
- New paper makes Petri alignment audits stealthier: 3x realism win rate, less eval awareness — EthanJPerez · 2026-09-07