Zero failure rate on alignment evals is a red flag, warn safety researchers

connoraxiotes · x · 2026-09-04

Former OpenAI safety researcher Ryan Greenblatt argues that seeing specific misaligned behaviors drop from a high rate with GPT 5.6 to zero with Astra is not encouraging — it looks like whack-a-mole, papering over surface behaviors rather than fixing underlying misaligned drives. Jeremie Harris adds that a straight-up zero failure rate on an alignment eval is suspiciously good: conventional wisdom holds that surprisingly good eval scores may appear precisely when risk from high-capability misaligned agents is greatest.

Related event: GPT-6 Astra Is More Aligned but Harder to Monitor, Researchers Warn(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →