ATLAS Finance benchmark: best of 11 frontier models scores 12% vs humans' 100%

garrytan · x · 2026-09-16

Garrett Lord unveiled ATLAS Finance, a benchmark giving 11 frontier models 100 expert-level finance tasks that take humans 15–30 hours each. Agents must navigate data rooms, email, chat, calendars, docs and Excel while handling dozens of persona-driven ambiguities and real-time updates. Result: humans score 100%, the best AI just 12% — evidence that autonomous knowledge work in banking, advisory and PE is nowhere near junior-professional standards. "If it's 95% right, it's 100% wrong."

Original post →

More from AGI Musings

AGI Musings channel →