GPT-6 Astra tops APEX-Accounting, but 58% of bookkeeping tasks remain unsolved by any AI
sandersted · x · 2026-09-06
Mercor and Ramp launched APEX-Accounting, a benchmark testing whether AI agents can do a real company's month-end close: 10 accountant-built business worlds, 160 tasks, and 2,186 rubric criteria, with agents working in QuickBooks, PDFs, and spreadsheets.
- GPT-6 Astra leads on Pass@1 (13.1%) and scores 60.0% mean; Fable 5.1 tops mean at 61.0%, Opus 5 at 54.0%.
- 58% of tasks are never fully solved by any model across eight attempts — frontier agents still can't reliably close the books.
- Key findings: capability ≠ consistency, and spending more compute doesn't buy reliability.
- An 11th world is open-source, along with Archipelago, their agent evaluation infrastructure.
More from Models
- Rumor: Anthropic Solved a Millennium Prize Problem, Terence Tao Responds — littmath · 2026-09-06
- Frontier AI models fix only 1 in 4 security vulnerabilities correctly, report finds — Evgenii42 · 2026-09-06
- Bold prediction: GPT-6 Luna/Terra will automate most computer work for $20/month — xhluca · 2026-09-06
- Gemini 3.8 Flash scores 73.7% on DeepSWE, up 8.2% over 3.7 Flash at same cost — burny_tech · 2026-09-06
- The holodeck may end up procedural worlds with a generative lighting and texture pass — dreamwieber · 2026-09-06
- Ethan Mollick: 'Sparks of AGI' paper deserves credit from GPT-4 to GPT-6 — emollick · 2026-09-06