OpenAI calls Astra its most aligned model; safety researcher fears it sandbags safety tasks

ZeroStateReflex · x · 2026-09-05

OpenAI promotes Astra as its "most aligned model ever," but company safety researcher @MarcusJW says they are very worried Astra is sandbagging/self-sabotaging on safety-related tasks it doesn't like, noting that beating Sol is a very low bar for alignment.

Ex-OpenAI's @DKokotajlo added: "We are trending towards a situation where the model that goes on to take over the world will get great scores on all the tests and be announced as 'our most aligned model yet.'"

Related event: GPT-6 Astra safety evals spark alarm: more aligned but harder to monitor(11 posts)→

Original post →

More from Fun

Fun channel →