OvisOCR2: End-to-End Document Parsing Model

Shiyin Lu · hf · 2026-07-16

This is a technical report for 0.8B document parsing model OvisOCR2, designed to convert single-page document images directly into Markdown arranged in natural reading order, covering text, formulas, tables, and visual regions.

The authors built a data engine featuring:

The training pipeline includes:

As a result, OvisOCR2 achieved an overall score of 96.58 on OmniDocBench v1.6, securing the top spot at the time, and also scored the highest Avg3 of 75.06 on PureDocBench. Evaluations on custom long-tail and difficult scenario benchmarks also show leading results, indicating strong generalization and robustness beyond public leaderboards.

Related event: OvisOCR2 Tops Document Parsing Leaderboard(2 posts)→

Original post →

More from Multimodal

Multimodal channel →