Simon Willison's pelican test shows GPT-6 Astra beats GPT-5.6 at every reasoning level
Simon Willison · rss · 2026-09-05
Simon Willison tested GPT-6 Astra with his classic pelican-riding-a-bicycle SVG benchmark across reasoning levels, comparing against GPT-5.6 Sol, Terra, and Luna.
- Astra clearly wins: even the best Sol output (xhigh, which he preferred over max) is abstract shapes, while every Astra level beats it; max is genuinely good.
- Flaw: below max, Astra still fails to reliably place pelican legs on both sides.
- Pricing: Astra costs 2x Sol ($10/M input, $50/M output vs $5/$30) but uses far fewer tokens per level; Astra low produces a better pelican than any Sol level for 9.55 cents.
- Curious detail: Astra and Luna both used 16 input tokens vs 26 for Sol/Terra, hinting they may be more closely related than OpenAI let on.
More from Models
- GPT-6 Astra Early Impressions: Reddit Users Call It the Most Capable Model Yet — imadade · 2026-09-05
- GPT-6 Astra's computer use wows users: clicks multiple micro buttons simultaneously — JasonBotterill · 2026-09-05
- OpenAI's early Astra rollout sparks claims it moved to cover up a discovered agent swarm — repligate · 2026-09-05
- New Artificial Analysis Scores Drop, But Qwen 3.8 27B Still Holds Up — RedditUsr2 · 2026-09-05
- Daily digest: Claude proves Fermat's Last Theorem in Lean, Apple's biggest launch wave — APPSO · 2026-09-05
- Plus Subscribers Angry: Astra Locked to Codex, Two Prompts Burn Entire 5-Hour Limit — Maximum-Face9536 · 2026-09-05