Tencent open-sources WeVisDoc: end-to-end document parsing model turns a page image into Markdown

xiaohu · x · 2026-09-19

Tencent's WeChat Vision team open-sourced WeVisDoc, an end-to-end document parsing model in 2B and 4B variants. One image in, one full-page output: Markdown body text, LaTeX formulas, HTML tables and restored reading order — all from a single VLM with no separate layout detector or OCR engine. The 2B suits resource-limited or clean-page cases; the 4B is stronger on photographed and degraded pages.

Related event: Tencent open-sources WeVisDoc, tops OmniDocBench at 95.38(3 posts)→

Original post →

More from Multimodal

Multimodal channel →