ExtractBench: Commercial VLM Recall Drops Below 35% on 50+ Page Enterprise Docs
llama_index · x · 2026-08-11
LlamaIndex has introduced ExtractBench, a comprehensive benchmark designed to evaluate information extraction from complex enterprise documents.
- Scale: Tested 14 systems across 370 enterprise documents, totaling 4,869 pages and 67 document types.
- Systems Evaluated: Included frontier VLMs, coding agents, and specialized extraction APIs.
- Evaluation Method: Fully deterministic metrics were used, avoiding LLM-as-a-judge.
- Key Finding: Beyond 50 pages, commercial VLMs collapse below 35% recall. While precision remains high, the models silently drop a significant portion of table rows.
More from Models
- xAI's Grok Build Lets Users Opt In to Share Coding Data for Model Training — XFreeze · 2026-08-11
- UnslothAI Confirms Its Acceleration Tools Work Well with Apple's MLX Framework — danielhanchen · 2026-08-11
- Open Source AI Popularity Leaderboard — aviaviaviavi · 2026-08-11
- Anthropic Criticized for AI Watermark Strategy That Could Drive Users Away — Brian821 · 2026-08-11
- Daybreak Blue in Codex Clarified: Not GPT-5.6 Cyber, But Security-Tailored GPT-5.6 Sol — Angaisb_ · 2026-08-11
- Gemini Criticized for Poor Performance Due to Legacy DeepSeek-R1 Reasoning Style — scaling01 · 2026-08-11