APEX-Accounting Benchmark: 58% Tasks Unsolved, Claude Fable 5 Takes the Lead
EdwardSun0909 · x · 2026-08-01
Mercor released APEX-Accounting, a benchmark evaluating AI agents on professional accounting tasks. The test includes 10 business worlds and 160 month-end close tasks, requiring cross-application reasoning and handling incomplete records.
Results show that frontier models cannot yet reliably execute professional accounting work: 58% of tasks were never fully solved by any model across eight attempts. Anthropic's Fable 5 led the board by meeting only 56.4% of grading criteria, followed by Meta Muse Spark 1.1 (52.6%) and OpenAI GPT-5.6 (51.5%). Furthermore, models occasionally solved a task but almost never maintained consistency across multiple runs.
More from Models
- DeepSeek's Suspected V4-Flash Model Endpoint Surfaces on Hugging Face — victormustar · 2026-08-01
- Elon Musk Announces Grok 4.5: Beats GPT-5.6 in Benchmarks, Launches CLI Coding Agent — elonmusk · 2026-08-01
- OpenAI Slashes Model Costs by 80%: Tech Breakthrough or Price War? — Tight-Grocery9053 · 2026-08-01
- Kimi K3 DSpark Upgrade: 1M Context Without Performance Degradation, 140k Downloads — BanghuaZ · 2026-08-01
- Exploring Claude Opus's Odd Visual Outputs with the "Dario and Amanda" Prompt — chicametipo · 2026-08-01
- Open-Source Pressure: Meme Jokes Vendor Cut Prices 80% Due to DeepSeek — InternationalGap3698 · 2026-08-01