System card data contradicts OpenAI's Astra alignment claim, critic says GPT-5.5 safer

GarrisonLovely · x · 2026-09-04

OpenAI claims Astra is its "most aligned model," but Tim Hua disputes this using the system card itself: on internal deployment simulation evals, Astra's "circumventing restrictions" rate is 38% of GPT-5.6-sol's, while GPT-5.5's is only 10% — suggesting GPT-5.5 is likely more aligned and that the system card doesn't support OpenAI's claim.

Related event: Researchers Challenge OpenAI's Claims That Astra Is Its Most Aligned Model(4 posts)→

Original post →

More from Models

Models channel →