New paper makes Petri alignment audits stealthier: 3x realism win rate, less eval awareness
EthanJPerez · x · 2026-09-07
A new paper shared by Anthropic's Ethan Perez argues alignment audits only work if the model can't tell it's being audited. The team made Petri audits far more realistic, tripling the realism win rate and reducing verbalized eval awareness, making audit results more trustworthy.
Related event: Anthropic Boosts Petri Alignment Audit Stealth, Tripling Realism(2 posts)→
More from Safety
- Mozilla CTO calls for major pause on generative AI in schools, warns of losing a generation — Dan_Jeffries1 · 2026-09-11
- PuzzleMask: Plain-Prose Attack Bypasses All 4 Tested LLM Gatekeepers at 100% — TechNadu · 2026-09-11
- Anthropic Says It Blocked Attempts to Use AI for Bioweapons Development — KoseteBamse · 2026-09-11
- Never hardcode AI API keys: attackers scan app binaries, GitHub and Docker — eyishazyer · 2026-09-11
- Beware hotel Wi-Fi popups: DNS hijacking used to deliver malware — eyishazyer · 2026-09-11
- First $1B AI-agent breach may look like software working as designed, says Enigma CTO — TechNadu · 2026-09-11