MonkeyOCRv2 shows document visual encoders can decide 13.2 points of accuracy
机器之心 · wechat · 2026-07-26
Why document visual encoders matter
MonkeyOCRv2 runs a controlled study where the language model, training data, optimization, and decoding are all fixed, and only the visual encoder changes. On 8 document-understanding benchmarks, MonkeyOCRv2-B reaches 57.2, beating the strongest baseline, OpenVision-B, by 13.2 points.
What the paper argues
- Document models need to preserve fine-grained evidence: digits, punctuation, formulas, layout, and reading order.
- The paper frames reconstruction as a test of “visual memory”: if a visual representation can help recover the page, it likely kept more useful evidence.
- Compared with plain text generation, adding reconstruction improves robustness when language shortcuts are weakened.
Key experiments
- With scrambled text and low resolution, recognition without reconstruction scores 55.4%, while reconstruction lifts it to 72.1%.
- On CHAOS-Bench, which pits visual evidence against language priors, the reconstruction variant improves page-level recall from 12.1 to 14.7 in one ablation.
- Replacing encoders with MonkeyOCRv2 improves downstream tasks including OCR, formula recognition, text detection, document tamper detection, and overlapping-text segmentation.
Broader result
The system reaches 83.3 on MDPBench across 17 languages, and the paper also introduces MonkeyDocv2, a public dataset of 113 million document images designed to make future comparisons more reproducible.
Related event: MonkeyOCRv2 Beats 3B Models with Sub-1B Parameters(2 posts)→
More from Multimodal
- Reddit clip shows an AI-generated aerobics video with stylized workout characters — fox-and-grapes6732 · 2026-07-26
- Fable turns a Cezanne-style city builder into a playable AI gesture game — emollick · 2026-07-26
- Opus 5 reportedly delivers remarkable 3D results, despite not winning everywhere — TAbrodi · 2026-07-26
- Yokohama street scene reconstructed with 1,018 photos and 4M 3DGS splats — janusch_patas · 2026-07-26
- AI image tools are getting good at turning composite photos into new compositions — RachelVT42 · 2026-07-26
- Creator shares a new Midjourney style with full prompt settings — azed_ai · 2026-07-26