MonkeyOCRv2 shows document visual encoders can decide 13.2 points of accuracy

机器之心 · wechat · 2026-07-26

Why document visual encoders matter

MonkeyOCRv2 runs a controlled study where the language model, training data, optimization, and decoding are all fixed, and only the visual encoder changes. On 8 document-understanding benchmarks, MonkeyOCRv2-B reaches 57.2, beating the strongest baseline, OpenVision-B, by 13.2 points.

What the paper argues

Key experiments

Broader result

The system reaches 83.3 on MDPBench across 17 languages, and the paper also introduces MonkeyDocv2, a public dataset of 113 million document images designed to make future comparisons more reproducible.

Related event: MonkeyOCRv2 Beats 3B Models with Sub-1B Parameters(2 posts)→

Original post →

More from Multimodal

Multimodal channel →