Ramp's Accounting Bench: best model hits only 21% accuracy across 137 tasks
willcb · x · 2026-09-18
Ramp partnered with accounting professionals to launch Accounting Bench, a benchmark of 137 tasks with grading rubrics grounded in everyday accounting workflows.
Even with three attempts, the best model achieved only 21% accuracy, showing that reliable agentic accounting remains an open challenge.
More from Research
- New academic spam: single-authored papers cold-pitching ARR service contributions — anmarasovic · 2026-09-18
- Index pretraining lifts humanoid zero-shot success from 8% to 56% — coreylynch · 2026-09-18
- A 'Life Diary' Eval Could Be the Toughest Test Yet for Continual Learning in LLMs — JohnnyNi13 · 2026-09-18
- Pretraining on Human Behavior Reportedly Boosts Task Success from 9% to 56% — Dr_Singularity · 2026-09-18
- Index pretraining lifts Helix 2.5 zero-shot success from 8% to 56%, generating 50 min of data per second — coreylynch · 2026-09-18
- LLMs got good at text and stayed bad at tables — and it's not just a training-data problem — FamiliarSlide7685 · 2026-09-18