MonkeyOCRv2 Sets Open-Source SOTA, Beating Qwen3-VL-235B

jiqizhixin · x · 2026-08-07

Huazhong University of Science and Technology and Kingsoft Office introduced MonkeyOCRv2, a vision-text foundation model for Document AI. It was trained on 113 million document images across 17 languages, learning to understand text and character strokes at the pixel level.

Experiments show its recognition accuracy improved from 58.7% to 67.3%. It sets a new open-source SOTA on MDPBench and outperforms much larger models like Qwen3-VL-235B in document parsing and understanding.

Original post →

More from Research

Research channel →