MonkeyOCRv2 Sets Open-Source SOTA, Beating Qwen3-VL-235B
jiqizhixin · x · 2026-08-07
Huazhong University of Science and Technology and Kingsoft Office introduced MonkeyOCRv2, a vision-text foundation model for Document AI. It was trained on 113 million document images across 17 languages, learning to understand text and character strokes at the pixel level.
Experiments show its recognition accuracy improved from 58.7% to 67.3%. It sets a new open-source SOTA on MDPBench and outperforms much larger models like Qwen3-VL-235B in document parsing and understanding.
More from Research
- 4B Open-Source Model Post-Trained with Castform Matches GPT-5.6 at 100x Lower Cost — petrusenko_max · 2026-08-07
- Researchers Propose 'Intelligent Labs' to Accelerate Scientific Discovery — MengdiWang10 · 2026-08-07
- Gary Marcus Debates Definition of Neurosymbolic AI: Tool Calls Don't Count — GaryMarcus · 2026-08-07
- GraphRAG in Real Estate: Layered Knowledge Graphs Reshape Property Management — adnan_hashmi · 2026-08-07
- Thinking Without Words: What Is Latent Reasoning in AI? — mikeflache · 2026-08-07
- GPT Codex Multi-Agent System Seals Over 15,000 Theorems in Lean — GiorgioPatrini · 2026-08-07