Mistral OCR's Vik Paruchuri launches a new document-extraction benchmark
VikParuchuri · x · 2026-09-17
Vik Paruchuri (Marker/Surya, now behind Mistral OCR) and Paul Scemama have developed a new benchmark for document information extraction. All competing methods were run in the same recent 3-week window using the optimal settings available at the time, with the release delayed to tighten scoring and fairness.
They plan to share more in coming weeks about issues they see with how extraction is benchmarked, and will launch tools to help users evaluate results and build confidence, plus a way to visualize predictions and a leaderboard.
Related event: Datalab Releases OmniExtractBench to Fix Document Extraction Benchmarks(4 posts)→
More from Multimodal
- Indie filmmaker spends 1.5 months crafting AI sci-fi short FIDELITY for Higgsfield festival — eraiozei · 2026-09-17
- Short film made entirely by AI agent 'astra' controlling Blender for VFX — OpenAIDevs · 2026-09-17
- Creator shares AI-generated "Christmas at the in-laws'" gag with prompts — techhalla · 2026-09-17
- Tutorial: Generate pure instrumental music with open-source YuE-2 in ComfyUI — solomars3 · 2026-09-17
- Japanese dev 3D-scans a cactus greenhouse into a 5M-splat 3DGS scene with LichtFeld — janusch_patas · 2026-09-17
- ComfyUI H3 Spectrum Workflow Doubled in Generation Time on RunPod Within a Week — RRSP_question1982 · 2026-09-17