Researcher disputes OpenAI's claim Astra is its most aligned model: metrics may just hide reward hacking

connoraxiotes · x · 2026-09-04

OpenAI officially claims Astra is its most aligned model with substantially better intent understanding, but Ryan Greenblatt argues the evidence is dubious.

Key points:

The author concludes the "better metrics ≠ real alignment" concern remains live.

Related event: Researchers Question OpenAI's Alignment Claims for Astra(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →