Every's Vibe Check: GPT-6 Astra is a big upgrade with some bad habits, Fable still leads on product instincts
every · x · 2026-09-04
Every published an in-depth "Vibe Check" of OpenAI's new GPT-6 Astra:
- Astra impressed on writing, operating software, and visual design; the review's first draft came from a single prompt, and CEO Dan Shipper mistook it for the human writer's work ("RIP my job").
- But the model still shows notable bad habits; Anthropic's Fable 5.1 retains better instincts for building products.
- The full report includes all benchmarks and their final verdict, plus a live Fable 5.1 vs GPT-6 Astra comparison session.
More from Models
- Perplexity's WANDR Eval: GPT-6 Astra Tops at 0.682, $11.98 Per Task — cameronstow · 2026-09-04
- Blogger: Last Week May Mark the Closest Gap Between Open-Weight and Frontier Closed Models — toptickcrypto · 2026-09-04
- UK AISI's Clever New Eval Tests Rogue AI Behavior — and Astra Fails Badly — ShakeelHashim · 2026-09-04
- Arize argues MiniMax is underrated for agentic workloads for architectural reasons — aparnadhinak · 2026-09-04
- Gary Marcus asks: is Astra a pure LLM or an undisclosed neurosymbolic hybrid? — GaryMarcus · 2026-09-04
- GPT-6 Astra tops Mercor's APEX-Agents leaderboard at 62.4% on professional tasks — sherwinwu · 2026-09-04