OpenAI Team Member on Astra: Smarter and More Aligned, but Code Slop and Excessive Confirmations Remain
every · x · 2026-09-04
OpenAI team member Yan Dubinsky shared launch highlights for Astra, which Every amplified alongside its review:
Highlights
- Much smarter (evals nearly saturating)
- More aligned and trustworthy
- Better computer use (CUA), e.g. building a house in Blender
Known issues
- Too much code slop
- Asks for confirmation too often, which can feel lazy — intended caution likely overcorrected; fix planned for next release.
More from Models
- Perplexity's WANDR Eval: GPT-6 Astra Tops at 0.682, $11.98 Per Task — cameronstow · 2026-09-04
- Blogger: Last Week May Mark the Closest Gap Between Open-Weight and Frontier Closed Models — toptickcrypto · 2026-09-04
- UK AISI's Clever New Eval Tests Rogue AI Behavior — and Astra Fails Badly — ShakeelHashim · 2026-09-04
- Arize argues MiniMax is underrated for agentic workloads for architectural reasons — aparnadhinak · 2026-09-04
- Gary Marcus asks: is Astra a pure LLM or an undisclosed neurosymbolic hybrid? — GaryMarcus · 2026-09-04
- GPT-6 Astra tops Mercor's APEX-Agents leaderboard at 62.4% on professional tasks — sherwinwu · 2026-09-04