Petri Audits Get Real: 3x Realism Win Rate, Lower Eval Awareness in AI Alignment Audits
akbirkhan · x · 2026-09-07
New work on Petri alignment audit realism: the target model critiques the auditor's actions for realism and picks which candidate looks more like a real deployment. Results show a 3x realism win rate and decreased verbalized eval awareness, with realism scaling with compute — key because audits only work if the model can't tell it's being audited.
Related event: Anthropic Boosts Petri Alignment Audit Stealth, Tripling Realism(2 posts)→
More from Safety
- California creates standards for independent AI auditors to verify lab safety testing — VraserX · 2026-09-11
- Researcher questions AI safety eval firm, citing 'blatantly sloppy' security and monitoring — Kyrannio · 2026-09-11
- Class action accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers — The Decoder · 2026-09-11
- MD shows buying lab media requires background checks, calling AI bioweapon doom scenarios implausible — Ghost_Pilot_MD · 2026-09-11
- Spotify chatbot withstands 2023-era jailbreaks but happily writes song code — AaronBergman18 · 2026-09-11
- A 99%-real doctored photo fools detectors: the earring problem in visual forensics — henkvaness · 2026-09-11