Marker author builds new document extraction benchmark, fixing major scoring flaws in existing ones

VikParuchuri · x · 2026-09-17

VikParuchuri (author of the Marker document parsing tool) released a new evaluation method for document extraction, pooling documents from ExtractBench (LlamaIndex), LongExtractionBench (Reducto), LongArray (Extend), and their own benchmark, with consistent rules for null fields and row matching.\n\nThey found dozens of significant errors in existing benchmarks — e.g., the same number of table mistakes can score either 100% or 0% on LongExtractionBench. Code and data are open-sourced.

Related event: Datalab Releases OmniExtractBench to Fix Document Extraction Benchmarks(4 posts)→

Original post →

More from Research

Research channel →