Unlimited-OCR reads 100-page PDFs in one shot
Shruti_0810 · x · 2026-07-20
China has open-sourced **Unlimited-OCR**, a 3B model for reading full documents in one pass. - Uses a **32K context window** to process an entire document at once instead of page by page - Preserves **cross-page context**, which helps with tables, references, and long-form documents - Reports **93%** on standard OCR benchmarks, about **+6** over the baseline - Claims **<0.11 error rate** beyond 40 pages - Runs **fully locally** and supports **Transformers, vLLM, Ollama, Docker, llama.cpp**, and more The post argues this approach makes enterprise OCR more reliable while eliminating per-page cloud OCR costs.
Related event: Baidu Open-Sources Unlimited-OCR for Long Documents(4 posts)→
More from Multimodal
- OpenArt AI demos a Video Remix tool that can transform an existing video — eyishazyer · 2026-07-21
- ElevenLabs raises ElevenMusic free usage to 400 tracks a month — lukeharries · 2026-07-21
- Google Gemini now watermarks every AI video it generates — Sure_Belt9076 · 2026-07-21
- GPT Image 2 Prompt Turns Product Shots into Surreal Reality-Bending Ads — aziz4ai · 2026-07-21
- LTX 2.3 LoRA demo changes a video’s camera angle — CQDSN · 2026-07-21
- Reddit user shares a surreal ChatGPT-generated poster — Creamy-Sundae-9991 · 2026-07-21