Tencent open-sources WeVisDoc: end-to-end document parsing model turns a page image into Markdown
xiaohu · x · 2026-09-19
Tencent's WeChat Vision team open-sourced WeVisDoc, an end-to-end document parsing model in 2B and 4B variants. One image in, one full-page output: Markdown body text, LaTeX formulas, HTML tables and restored reading order — all from a single VLM with no separate layout detector or OCR engine. The 2B suits resource-limited or clean-page cases; the 4B is stronger on photographed and degraded pages.
Related event: Tencent open-sources WeVisDoc, tops OmniDocBench at 95.38(3 posts)→
More from Multimodal
- ComfyMax Music Video Director: an AI music video workflow in progress — DanielVeres · 2026-09-19
- Jina AI's jina-ocr-v1 document understanding model trends on Hugging Face — jinaai · 2026-09-19
- Reddit user seeks an autonomous agent that builds and tunes ComfyUI video workflows end-to-end — Hopeful-Election-783 · 2026-09-19
- Single-Word Prompt Series: Aefauld, a Scots Word for Sincerity, in Midjourney — tisch_eins · 2026-09-19
- Awwwards-mcp turns award-winning sites into an agent-readable design library — _insane7 · 2026-09-19
- Story Illustrator: Open-Source Tool Auto-Illustrates Stories via Local LLM + ComfyUI — Natrimo · 2026-09-19