Datalab Releases OmniParseBench, an Open OCR Benchmark It Doesn't Top
Datalab (Vik Paruchuri's team) has released a public OCR benchmark, OmniParseBench, featuring 16288 pass/fail tests with both data and scoring code fully open-sourced — and notably disclosed that its own product does not top the leaderboard, a rare move in an era of self-promotional benchmark releases.
Confirmed
- Benchmark scale: 16288 pass/fail tests drawn from 2343 documents and 2937 pages, covering 92 languages, with unified and interpretable scoring.
- Coverage: handwriting, tables, multilingual text, math formulas, forms, and other hard OCR edge cases.
- Motivation: Vik Paruchuri said the team has used olmOCR-bench for over a year and considers it the fairest, most consistent benchmark in the field, but it is now saturated — remaining gains mostly come from matching its output format rather than genuine improvements in reading ability.
- The team has committed to continually updating test cases to track progress in the field.
Why it matters
- Pass/fail scoring is more auditable than fuzzy scoring, and open-sourced data and code aid reproducibility.
- Coverage of long-standing pain points like handwriting and formulas makes it a lasting yardstick for multilingual OCR progress.
- The publisher's disclosure that its own product is not number one boosts the benchmark's neutrality and credibility.
2026-10-10 ~ 2026-10-10 · 5 related posts
Primary sources
- Datalab launches OmniParseBench, a fair 16k-test OCR benchmark — and admits it's not No.1 — VikParuchuri ·
- olmOCR-bench is saturated, so Datalab built auditable OmniParseBench: 16,288 tests, 92 languages — VikParuchuri ·
- OmniParseBench Debuts: New OCR Benchmark for Handwriting, Tables, Math, Forms — VikParuchuri ·
- [source] Datalab launches OmniParseBench, a fair 16k-test OCR benchmark — and admits it's not No.1 — VikParuchuri · 2026-10-10
- [source] olmOCR-bench is saturated, so Datalab built auditable OmniParseBench: 16,288 tests, 92 languages — VikParuchuri · 2026-10-10
- olmOCR-bench has saturated — gains now come from format matching, not better reading — VikParuchuri · 2026-10-10
- OmniParseBench tracks OCR edge cases: handwriting, tables, math, forms — VikParuchuri · 2026-10-10
- [source] OmniParseBench Debuts: New OCR Benchmark for Handwriting, Tables, Math, Forms — VikParuchuri · 2026-10-10