Datalab launches OmniExtractBench: 620 docs to fix biased structured-extraction benchmarks
VikParuchuri · x · 2026-09-17
Structured extraction is hard to eval, and existing benchmarks are biased or have flawed scoring/ground truth. Vik Paruchuri's Datalab team built OmniExtractBench:
- Composition: 620 documents drawn from ExtractBench (LlamaIndex), LongExtractionBench (Reducto), LongArray (Extend), and their own benchmark.
- Fixes: consistent rules for null fields and row matching; they flag that in LongExtractionBench the same number of table mistakes can score either 100% or 0%.
- Results: Datalab and Reducto top the leaderboard.
- Code and data are publicly available.
Related event: Datalab Releases OmniExtractBench to Fix Document Extraction Benchmarks(4 posts)→
More from Research
- Mozilla's 91-page report: open-weight AI now only ~4 months behind the frontier — rohanpaul_ai · 2026-09-17
- Protein Models ESMC and ESMFold2 Land on HuggingFace with NVIDIA-Built Fused Triton Kernels — AllThingsApx · 2026-09-17
- Panel Discussion on AI in Mathematical Research from Sept 15 Comes Highly Recommended — MindlessPapaya8463 · 2026-09-17
- OpenAI hints at millennium problem breakthrough ahead of Sept 29 DevDay — haider1 · 2026-09-17
- Vitalik: cybersecurity favors defense in the AI-hacking era, thanks to formal verification — jessi_cata · 2026-09-17
- PufferLib 5.0 hits 60M steps/sec single-GPU RL training, solves Breakout in under a second — jsuarez · 2026-09-17