Baidu Open-Sources Long-Doc OCR Model
brianrkelly · x · 2026-07-20
Baidu has open-sourced an OCR model named **Unlimited-OCR**. With only 3 billion parameters, it is designed to read entire long documents in one go, claiming the ability to directly parse 100-page PDFs. It supports 32K context, multiple languages, and can run on local hardware.\n\nThe post also shares several results: it achieves 93% on standard parsing benchmarks, 6 points higher than the baseline, and for documents over 40 pages, the error rate drops below 0.11. The model is compatible with Transformers, vLLM, SGLang, Docker, Ollama, and llama.cpp. The author notes it has reached 1.9 million downloads on Hugging Face and was developed to push the boundaries of DeepSeek-OCR even further.
Related event: Baidu Open-Sources Unlimited-OCR for Long Documents(4 posts)→
More from Models
- Epoch AI Live Streams GPT-5.6 Playing Slay the Spire — dr_cintas · 2026-07-21
- Side-by-side model test lands both answers on the first try, then shifts to loop engineering — glenbeer · 2026-07-21
- Emad Mostaque says Kimi K3 inference costs could fall 10x to 50x soon — rohanpaul_ai · 2026-07-21
- Reddit asks whether Kimi K3 is already good enough for production agents — CommercialClient2408 · 2026-07-21
- Korean startup says its model scored 44 on AAII and matches DeepSeek V4 Pro — JungWooHa2 · 2026-07-21
- OpenAI’s GPT-6 is predicted to be far more efficient than Fable — bindureddy · 2026-07-21