MonkeyOCRv2: Document AI Vision Model
VLRLab-OCR · hf · 2026-07-15
This paper introduces MonkeyOCRv2, a vision-text foundation model designed for Document AI.
Key highlights include:
- The authors constructed MonkeyDoc v2, a pre-training corpus consisting of 113 million document images across 17 languages.
- Pre-training is conducted simultaneously on image-to-text generation and pixel-level document reconstruction to preserve both character shape and layout information.
- Delivers consistent improvements across 5 tasks: text recognition, formula recognition, text detection, document tamper detection, and overlapping text segmentation.
- By freezing the encoder, it forms a 0.7B document parsing model that sets a new open-source SOTA on MDPBench, achieving an absolute improvement of 2.8% over the previous best 3B dots.mocr, while having a visual encoder approximately 11 times smaller.
- The same encoder also outperforms CLIP, DINO, and SAM as a visual frontend for document understanding models.
Related event: MonkeyOCRv2: A New Document AI Foundation Model(2 posts)→
More from Multimodal
- Qwen3-VL flags a burned-in clinic name that another PII model missed — MaziyarPanahi · 2026-07-22
- OpenAI Build Week demo shows Better Backgrounds for video calls — cjami · 2026-07-22
- Open-source Image SDK unifies 8 providers behind one API — Ok_Window_2596 · 2026-07-22
- Snap Spectacles demo turns World Cup tracking data into editable photoreal replay — stspanho · 2026-07-22
- NKD VFX Tools turns AI image generation into a controllable VFX pipeline — Nekodificador · 2026-07-22
- Midjourney image made from a single character reference — gen_ericai · 2026-07-22