OpenAI Astra safety data: more capable model, zero misaligned cyber attacks vs Sol's 56%

VoidStateKate · x · 2026-09-02

The author digs into the most interesting part of OpenAI's Astra announcement:

The key insight: "can the model do something dangerous?" and "will it decide to when it shouldn't?" are different problems — and for once, capability rose sharply while tested misaligned behavior dropped. Either stronger models naturally become more aligned, or they're scheming somewhere we can't see.

Original post →

More from Models

Models channel →