Baidu Open-Sources Long-Document OCR Model
TheTimHayden · x · 2026-07-20
Baidu has open-sourced Unlimited-OCR, an OCR model designed for entire documents, focusing on processing long documents in one go rather than splitting them by page.
It has 3 billion parameters but only activates 500 million during inference. It supports local execution, features a 32K context window, and better preserves cross-page text, formulas, tables, and reading order, directly outputting structured Markdown.
Data provided by the author includes: 93% accuracy on standard benchmarks, 6 points higher than the baseline, with an error rate remaining below 0.11 for documents over 40 pages. It also supports multiple languages, achieving 2.12 million monthly downloads on Hugging Face and 14,600 stars on GitHub.
Related event: Baidu Open-Sources Unlimited-OCR for Long Documents(4 posts)→
More from Multimodal
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11
- MiniMax Music Production Toolkit 2.5 for ComfyUI adds full mastering chain — Vivid_Promise1700 · 2026-09-11
- New Node Finder for ComfyUI ranks fresh nodes by star velocity and recency — Luke2642 · 2026-09-11
- Using a finisher move on one mosquito with MiniMax H3 MAX — the bug survives — Hailuo_AI · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11