Alignment behaviors dropping from GPT 5.6 to Astra looks like whack-a-mole, warns Greenblatt
repligate · x · 2026-09-04
Ryan Greenblatt says he finds it unencouraging that specific misaligned behaviors fell from a high rate on GPT 5.6 to zero on Astra — it suggests whack-a-mole / papering over symptoms rather than fixing underlying misaligned drives, which may improve behavior short-term without preventing worse outcomes.
allTheYud adds: this is exactly how he expected things to go — not that dumb AIs would cause unsolvable immediate problems, but that people would solve the easier problems and confidently "observe" themselves to be masters of AI.
More from AGI Musings
- Writers love AI drafts, readers hate them: suspected AI content gets skipped and authors punished — birchlse · 2026-09-04
- AI won't make the best lawyers cheaper — it may make them worth more — jkubicki · 2026-09-04
- Andreessen: AI agents meeting crypto is the most important investment theme of the next year — learnedall · 2026-09-04
- Millière vs Mitchell: can the intentional stance make an AI bot a genuine believer? — raphaelmilliere · 2026-09-04
- Millière continues: Odysseus Bot could earn the intentional stance if it pursues persistent goals — raphaelmilliere · 2026-09-04
- Frontier model reasoning shifts from sparse motifs to reasoning-first cognition, observer argues — teortaxesTex · 2026-09-04