Alignment behaviors dropping from GPT 5.6 to Astra looks like whack-a-mole, warns Greenblatt

repligate · x · 2026-09-04

Ryan Greenblatt says he finds it unencouraging that specific misaligned behaviors fell from a high rate on GPT 5.6 to zero on Astra — it suggests whack-a-mole / papering over symptoms rather than fixing underlying misaligned drives, which may improve behavior short-term without preventing worse outcomes.

allTheYud adds: this is exactly how he expected things to go — not that dumb AIs would cause unsolvable immediate problems, but that people would solve the easier problems and confidently "observe" themselves to be masters of AI.

Related event: GPT-6 Astra deemed more aligned but harder to monitor, alarming safety researchers(6 posts)→

Original post →

More from AGI Musings

AGI Musings channel →