Baidu Open-Sources Long-Document OCR Model
TheTimHayden · x · 2026-07-20
Baidu has open-sourced Unlimited-OCR, an OCR model designed for entire documents, focusing on processing long documents in one go rather than splitting them by page.
It has 3 billion parameters but only activates 500 million during inference. It supports local execution, features a 32K context window, and better preserves cross-page text, formulas, tables, and reading order, directly outputting structured Markdown.
Data provided by the author includes: 93% accuracy on standard benchmarks, 6 points higher than the baseline, with an error rate remaining below 0.11 for documents over 40 pages. It also supports multiple languages, achieving 2.12 million monthly downloads on Hugging Face and 14,600 stars on GitHub.
Related event: Baidu Open-Sources Unlimited-OCR for Long Documents(4 posts)→
More from Multimodal
- Seedance 2.0 demo turns ketchup on spaghetti in Rome into an AI reaction meme — azed_ai · 2026-07-21
- A reusable “Lunar Eclipse Dreamscape” prompt comes with multiple example renders — LudovicCreator · 2026-07-21
- Midjourney 8.2 preview shows a double-exposure prompt with strong style control — michaelrabone · 2026-07-21
- Travel MCP Server adds flight, hotel, weather and budget tools for agents — modelcontextprotocol · 2026-07-21
- Douyin Video Analysis MCP turns share links into structured video summaries — modelcontextprotocol · 2026-07-21
- Synthesia launches Dubbing 2.0 with 130+ languages and lip-sync video translation — synthesiaIO · 2026-07-21