GPT-6 Astra system card draws fire: OpenAI claims 'most aligned model' ever
TheZvi · x · 2026-09-09
Zvi analyzes OpenAI's GPT-6 Astra system card, where OpenAI claims Astra is 'the most intelligent and most aligned model available' — not just within OpenAI, but globally.
- The piece questions what 'most aligned' even means: how is alignment measured, and why would it beat Claude Fable 5.1?
- Highlights a pointed exchange: keltan asks 'how tf are you measuring alignment? Being able to measure that would save the world'; OpenAI's roon answers 'low rates of cheating,' which Rob Miles corrects to 'detected cheating.
- The author considers Astra clearly an excellent model, but warns the bold claims and severe monitorability problems risk souring the release.
More from AGI Musings
- Musk: AI and robots will more than double the global economy in under 10 years — elonmusk · 2026-09-09
- Pretraining Researcher Resigns From Anthropic, Warns of Race to Self-Improving Superintelligence — DevDminGod · 2026-09-09
- Ben Thompson's 'Write Things Down': Why GTD-Style Externalization Works for Humans and AI — timigod · 2026-09-09
- Hugging Face Co-founder Thom Wolf Argues AI Math Isn't Solved: Counterexamples Over Full Proofs — Thom_Wolf · 2026-09-09
- We already got AGI a while ago — goal-post shifting masked it, argues Reddit post — Insane_Artist · 2026-09-09
- Prediction: generative AI will add "layers" over reality, not just rebuild games — danfaggella · 2026-09-09