olmOCR-bench is saturated, so Datalab built auditable OmniParseBench: 16,288 tests, 92 languages
VikParuchuri · x · 2026-10-10
Datalab released OmniParseBench: 16,288 pass/fail tests on 2,937 pages from 2,343 documents across 92 languages, with open data and scorer code. It exists because olmOCR-bench has saturated — remaining gains reflect format matching, not reading — and ParseBench's per-capability metrics share no common unit. Every test is auditable, tagged, and explorable by slice.
Related event: Datalab Releases OmniParseBench, an Open OCR Benchmark It Doesn't Top(5 posts)→
More from Models
- Liquid AI's decision model d1 lands on Vercel AI Gateway with vision support — maximelabonne · 2026-10-10
- Open TTS Leaderboard adds Paradee-8M, a Kokoro-82M distill matching WER at 1/10 params — realmrfakename · 2026-10-10
- Kimi gateway latency test ranks GitHub first, Neon second, ngrok third — mariorod1 · 2026-10-10
- Grok in group chats is 'really nice', but users note messages are no longer private — Angaisb_ · 2026-10-10
- Grok bot buys on Amazon first try while shopping agent Muse fails twice — jamesperkins · 2026-10-10
- Why long AI chats get expensive: full-history resends and brittle prefix caching — ClickOk5811 · 2026-10-10