Baidu Open-Sources Unlimited-OCR for Long Documents
thisdudelikesAI · x · 2026-07-20
Baidu’s Unlimited-OCR has been open-sourced as a local document OCR model.
The post says it is a 3B model that can process long PDFs in a single pass with a 32K context window, supports multiple languages, and runs entirely on local hardware. It is claimed to reach 93% on a standard parsing benchmark, keep error rates below 0.11 after 40 pages, and work with Transformers, vLLM, SGLang, Docker, Ollama, and llama.cpp.
The comparison point is that cloud OCR services charge per page, while this setup avoids upload limits, per-page billing, and sending private documents to a remote server. The post also says it has surpassed 1.9 million downloads on Hugging Face.
Related event: Baidu Open-Sources Unlimited-OCR for Long Documents(4 posts)→
More from Infra
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11
- Can a 7900 XTX 24GB run Qwen locally? Reddit seeks ROCm tok/s benchmarks — thenomadexplorerlife · 2026-09-11
- RTK Terminal Compression Cuts Tokens but Leaves Your AI Coding Bill Unchanged — Bartaseth · 2026-09-11
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11