Doc Understanding Benchmark Shows Clear Gains

echen · x · 2026-07-13

Surge introduced their **GDP.pdf** public evaluation set, designed to test whether frontier models can understand "documents that keep the real world running," such as financial files, dosage tables, and compensation clauses. Using this public set, @reductoai conducted experiments showing that after plugging in their document parsing capabilities, model quality improved by an average of **9 percentage points**, while inference token consumption dropped by **13%**. The biggest gains were seen in dense engineering tasks like wiring diagrams and cross-referenced tables, with improvements ranging from **7% to 23%**. The author emphasized that this public set allows teams building document AI stacks to directly build on and compare against it, and welcomed further discussion.

Original post →

More from Apps

Apps channel →