GPT-6 Astra tops Zapier's AutomationBench at 41.4%, first model ever to clear 40%
sandersted · x · 2026-09-04
Zapier reports GPT-6 Astra scored 41.4% on AutomationBench — the highest ever and the first above 40% (previous best: GPT-5.6-Sol at 28.8%). The benchmark tests agents on end-to-end workflows with 47 real tools across Sales, Marketing, Operations, Support, Finance and HR, scored deterministically on final environment state. Astra excels at arithmetic across messy sources (reconciliation, deal review, vendor scorecards); Operations/Support are its strongest domains, HR weakest. Current leaderboard: Claude Fable 5.1 + Opus 5 fallback 31.4% ($2.45/task), Gemini 3.7 Flash 30.44% ($0.61).
More from coding & agent
- Grok Bot Launches for Enterprise, Free for Grok and Cursor Customers for Two Weeks — mattyp · 2026-09-04
- Perplexity says Astra excels at computer use, coming to Comet browser — AravSrinivas · 2026-09-04
- OpenAI showcases GPT-6 Astra: 3D London history, matcha site and DEF CON puzzle via parallel agents — OpenAI · 2026-09-04
- swyx Burned 20B Tokens Stress-Testing Astra on Real AI Engineering Tasks — All for Under $6/Hour — charliermarsh · 2026-09-04
- LangChain launches free LangSmith Essentials course covering the full agent dev lifecycle in 60 minutes — Hacubu · 2026-09-04
- OpenAI's Lukasz Kaiser: even we can't pinpoint what caused the Christmas coding-agent jump — a_karvonen · 2026-09-04