Baidu Open-Sources Long-Doc OCR Model

brianrkelly · x · 2026-07-20

Baidu has open-sourced an OCR model named **Unlimited-OCR**. With only 3 billion parameters, it is designed to read entire long documents in one go, claiming the ability to directly parse 100-page PDFs. It supports 32K context, multiple languages, and can run on local hardware.\n\nThe post also shares several results: it achieves 93% on standard parsing benchmarks, 6 points higher than the baseline, and for documents over 40 pages, the error rate drops below 0.11. The model is compatible with Transformers, vLLM, SGLang, Docker, Ollama, and llama.cpp. The author notes it has reached 1.9 million downloads on Hugging Face and was developed to push the boundaries of DeepSeek-OCR even further.

Related event: Baidu Open-Sources Unlimited-OCR for Long Documents(4 posts)→

Original post →

More from Models

Models channel →