Baidu Open-Sources Unlimited-OCR for Long Documents
thisdudelikesAI · x · 2026-07-20
Baidu’s Unlimited-OCR has been open-sourced as a local document OCR model.
The post says it is a 3B model that can process long PDFs in a single pass with a 32K context window, supports multiple languages, and runs entirely on local hardware. It is claimed to reach 93% on a standard parsing benchmark, keep error rates below 0.11 after 40 pages, and work with Transformers, vLLM, SGLang, Docker, Ollama, and llama.cpp.
The comparison point is that cloud OCR services charge per page, while this setup avoids upload limits, per-page billing, and sending private documents to a remote server. The post also says it has surpassed 1.9 million downloads on Hugging Face.
Related event: Baidu Open-Sources Unlimited-OCR for Long Documents(4 posts)→
More from Infra
- Gritt raises a new round to automate solar array installation and maintenance — rebeccakaden · 2026-07-21
- Refactoring 150k LOC Takes 96 Hours: Is Compute the Bottleneck for AI Coding? — robleclerc · 2026-07-21
- Seeking Recommendations: Essential Local Small Models (Audio/Vision/TTS) — DeepOrangeSky · 2026-07-21
- Compute Allocation Limits: The Root Cause of Missing Architecture Innovation in European LLMs — IgorCarron · 2026-07-21
- AI Energy Footprint Pales Compared to Transport and Agriculture — dreamwieber · 2026-07-21
- SkyPilot comes out of stealth with claims of 10x faster AI time-to-intelligence — songhan_mit · 2026-07-21