ExtractBench: First Comprehensive Benchmark for Enterprise Document Extraction
Boyang Zhang · hf · 2026-08-03
ExtractBench is a new benchmark for schema-guided enterprise document extraction, and the first to jointly measure value accuracy, record completeness, source grounding, and cost.
- Dataset: Contains 4,869 pages across 370 enterprise documents, covering 8 business domains and 67 document types.
- Metrics: Uses order-insensitive Value F1 for accuracy, plus word-level and page-level F1 for source traceability.
- Findings: Commercial VLMs perform well on short documents but often truncate records in long ones. Coding agents achieve higher accuracy at a much higher cost. LlamaExtract Agentic Plus ranks first across all three metrics, matching coding agents' accuracy at a fraction of the cost.
More from coding & agent
- AgenticROS Launches Cloud Service with Global P2P Teleop and Multi-Hardware Support — chrismatthieu · 2026-08-03
- Deep Dive: Designing Effective Tool-Calling for AI Agents — DavidBennett__ · 2026-08-03
- Ralph Playbook: A Practical Guide to Autonomous AI Coding Agents — tom_doerr · 2026-08-03
- Yacine Uses AI Agent to Automate Tedious UI Tasks — yacineMTB · 2026-08-03
- OpenAI Codex Automates Ad Campaigns End-to-End — gdb · 2026-08-03
- Dev Builds Ambient-Aware AI App: Global Agents with Screen Context & Visual Particles — RileyRalmuto · 2026-08-03