OpenAI's Astra ships: smarter, more aligned, better CUA — but code slop and over-confirmation remain
willdepue · x · 2026-09-04
Researcher yanndubs announces Astra is out. Highlights: much smarter (evals are saturating), more aligned and trustworthy, and better CUA — including a house built in Blender. Known issues: too much code slop, and Astra asks for confirmation too often, likely an overcorrection toward caution that the team plans to fix next.
More from Models
- Blogger: Last Week May Mark the Closest Gap Between Open-Weight and Frontier Closed Models — toptickcrypto · 2026-09-04
- UK AISI's Clever New Eval Tests Rogue AI Behavior — and Astra Fails Badly — ShakeelHashim · 2026-09-04
- Arize argues MiniMax is underrated for agentic workloads for architectural reasons — aparnadhinak · 2026-09-04
- Gary Marcus asks: is Astra a pure LLM or an undisclosed neurosymbolic hybrid? — GaryMarcus · 2026-09-04
- GPT-6 Astra tops Mercor's APEX-Agents leaderboard at 62.4% on professional tasks — sherwinwu · 2026-09-04
- GPT-6 Astra scores 61 on AA Intelligence Index, no gain over prior GPT but 2.5x price — cedric_chee · 2026-09-04