GPT-6 Astra is here: evals saturating, more trustworthy — and more code slop
yanndubs · x · 2026-09-04
A team member (Yannt Dubois) walks through GPT-6 Astra:
- Highlights: notably smarter (evals are saturating), more aligned and trustworthy, stronger computer-use ability — including a house Astra built in Blender
- Known issues: too much code slop; asks for confirmation too often, which reads as lazy — an overcorrection toward caution that the team plans to fix
- Caveat: Astra follows instructions even more strictly than 5.6, including unwanted ones buried in old skills — clean those up before using it
Related event: GPT-6 Astra Team Member Highlights Gains, Admits More Code Slop(4 posts)→
More from Models
- Every's Vibe Check: GPT-6 Astra Is a Big Upgrade, but Anthropic's Fable Still Has Better Product Instincts — every · 2026-09-04
- GPT-6 Astra Nukes ARC-AGI-3: Score Jumps from 8% to 63%, 98.6% with Adapter — haider1 · 2026-09-04
- Sam Altman Officially Launches GPT-6 Astra, Claiming Best-in-World Computer Use and Coding — eyishazyer · 2026-09-04
- OpenAI claims GPT-6 Astra SOTA on FrontierMath Tier 4, ARC-AGI 3, TerminalBench-4.0 — dair_ai · 2026-09-04
- Five releases in 48 hours: GPT-6 Astra, Fable 5.1, Gemini 3.8 Flash and more — dr_cintas · 2026-09-04
- Matthew Bellerman Tests GPT-6 Astra Early: 'The Best Model I've Ever Used, Period' — every · 2026-09-04