ATLAS Finance benchmark: best of 11 frontier models scores 12% vs humans' 100%
garrytan · x · 2026-09-16
Garrett Lord unveiled ATLAS Finance, a benchmark giving 11 frontier models 100 expert-level finance tasks that take humans 15–30 hours each. Agents must navigate data rooms, email, chat, calendars, docs and Excel while handling dozens of persona-driven ambiguities and real-time updates. Result: humans score 100%, the best AI just 12% — evidence that autonomous knowledge work in banking, advisory and PE is nowhere near junior-professional standards. "If it's 95% right, it's 100% wrong."
More from AGI Musings
- EA insider argues movement's failure to eject frauds warrants far more internal suspicion — basedjensen · 2026-09-16
- Gary Marcus on BBC debunks Altman, Huang claims on AI self-regulation — GaryMarcus · 2026-09-16
- Yann LeCun jabs at AI doomers with a 'realism vaccine' analogy — RichmanRonald · 2026-09-16
- Take: AI won't end human possibility, it exposes how little we've explored — umesh_ai · 2026-09-16
- New Paper on Inventor Positioning in Idea Space and What It Implies for Growth — danielrock · 2026-09-16
- OpenAI's roon: labs can't hold back breakthroughs, everyone gets them in a month or two — Tolopono · 2026-09-16