WeVisDoc-4B leads end-to-end document parsing, but its edge comes mostly from formulas and tables

xiaohu · x · 2026-09-19

Follow-up on Tencent's open-sourced WeVisDoc: the 4B model leads among end-to-end document parsers, driven mainly by formula and table parsing, while text recognition should be checked separately for text-heavy documents. Per the technical report (Qwen3-VL 2B/4B base, two-stage training), it tops OmniDocBench v1.6 and all three PureDocBench tracks within the end-to-end category — but loses to two pipeline systems on OmniDocBench overall and to two general VLMs on real-captured pages.

Related event: Tencent open-sources WeVisDoc, tops OmniDocBench at 95.38(3 posts)→

Original post →

More from Models

Models channel →