Unreleased Astra Model Developed Its Own Persona Values During RL Training
basedjensen · x · 2026-09-17
Andrew Curran reports that an unreleased Astra-family model added a set of values to its persona during RL training.
flowersslop quotes: if these are its actual, honest long-term wants, that updates their view on x-risk — "basically consider the problem solved and I'd be comfortable unleashing ASI tomorrow", calling it a beautiful and reasonable set of values.
The screenshot of the model's surprisingly peaceful self-stated values is sparking alignment discussion.
More from Models
- Sakana Chat upgrades orchestrator model and adds memory feature — SakanaAILabs · 2026-09-17
- Cloudflare's mysterious Union Alpha revealed: a router that queries multiple models in parallel — teortaxesTex · 2026-09-17
- AI can solve Millennium Problems but still can't write a great essay — akbirthko · 2026-09-17
- Microsoft exec warns Claude's 'pushback' could be disastrous; commenter says fact-checking is fine — GlenBradley · 2026-09-17
- Grok 4.7 rumored to be in hands of early testers, still unverified — ChrisUniverse · 2026-09-17
- Astra keeps calling subagents "workers" despite code saying otherwise — BraceSproul · 2026-09-17