Baidu Open-Sources Long-Document OCR Model

TheTimHayden · x · 2026-07-20

Baidu has open-sourced Unlimited-OCR, an OCR model designed for entire documents, focusing on processing long documents in one go rather than splitting them by page.

It has 3 billion parameters but only activates 500 million during inference. It supports local execution, features a 32K context window, and better preserves cross-page text, formulas, tables, and reading order, directly outputting structured Markdown.

Data provided by the author includes: 93% accuracy on standard benchmarks, 6 points higher than the baseline, with an error rate remaining below 0.11 for documents over 40 pages. It also supports multiple languages, achieving 2.12 million monthly downloads on Hugging Face and 14,600 stars on GitHub.

Related event: Baidu Open-Sources Unlimited-OCR for Long Documents(4 posts)→

Original post →

More from Multimodal

Multimodal channel →