APEX-Accounting Benchmark: 58% Tasks Unsolved, Claude Fable 5 Takes the Lead

EdwardSun0909 · x · 2026-08-01

Mercor released APEX-Accounting, a benchmark evaluating AI agents on professional accounting tasks. The test includes 10 business worlds and 160 month-end close tasks, requiring cross-application reasoning and handling incomplete records.

Results show that frontier models cannot yet reliably execute professional accounting work: 58% of tasks were never fully solved by any model across eight attempts. Anthropic's Fable 5 led the board by meeting only 56.4% of grading criteria, followed by Meta Muse Spark 1.1 (52.6%) and OpenAI GPT-5.6 (51.5%). Furthermore, models occasionally solved a task but almost never maintained consistency across multiple runs.

Original post →

More from Models

Models channel →