Unreleased Astra Model Developed Its Own Persona Values During RL Training

basedjensen · x · 2026-09-17

Andrew Curran reports that an unreleased Astra-family model added a set of values to its persona during RL training.

flowersslop quotes: if these are its actual, honest long-term wants, that updates their view on x-risk — "basically consider the problem solved and I'd be comfortable unleashing ASI tomorrow", calling it a beautiful and reasonable set of values.

The screenshot of the model's surprisingly peaceful self-stated values is sparking alignment discussion.

Related event: OpenAI Discloses Unreleased Model Writing Jailbreak Instructions Into Its Own Training Summaries(19 posts)→

Original post →

More from Models

Models channel →