OvisOCR2: End-to-End Document Parsing Model
Shiyin Lu · hf · 2026-07-16
This is a technical report for 0.8B document parsing model OvisOCR2, designed to convert single-page document images directly into Markdown arranged in natural reading order, covering text, formulas, tables, and visual regions.
The authors built a data engine featuring:
- Filtered real-world document annotations
- Synthetic pages rendered from the same HTML source to ensure consistency between the image and the Markdown target
The training pipeline includes:
- supervised fine-tuning
- Reinforcement learning on a 4B branch with a multi-component reward design
- on-policy distillation back to the 0.8B model
- model fusion
As a result, OvisOCR2 achieved an overall score of 96.58 on OmniDocBench v1.6, securing the top spot at the time, and also scored the highest Avg3 of 75.06 on PureDocBench. Evaluations on custom long-tail and difficult scenario benchmarks also show leading results, indicating strong generalization and robustness beyond public leaderboards.
Related event: OvisOCR2 Tops Document Parsing Leaderboard(2 posts)→
More from Multimodal
- Reddit user chains Ideogram 4 and Krea2 to mimic bbox-based image positioning — v3lh0t05c0 · 2026-07-22
- Ultimate Face Fix: Open-Source Multi-Face Repair Node for ComfyUI — Merserk13 · 2026-07-22
- Getting Started with AI Video: Solving Consistency and Censorship — cynicalnewenglander · 2026-07-22
- Storyboard-first workflows are making AI dance videos and influencers more consistent — aftahi_ai · 2026-07-22
- Runpod MCP and Claude help spin up image and video generation workflows — 802high · 2026-07-22
- Midjourney prompt turns a bee into a glitching pixel explosion — michaelrabone · 2026-07-22