WeVisDoc-4B leads end-to-end document parsing, but its edge comes mostly from formulas and tables
xiaohu · x · 2026-09-19
Follow-up on Tencent's open-sourced WeVisDoc: the 4B model leads among end-to-end document parsers, driven mainly by formula and table parsing, while text recognition should be checked separately for text-heavy documents. Per the technical report (Qwen3-VL 2B/4B base, two-stage training), it tops OmniDocBench v1.6 and all three PureDocBench tracks within the end-to-end category — but loses to two pipeline systems on OmniDocBench overall and to two general VLMs on real-captured pages.
Related event: Tencent open-sources WeVisDoc, tops OmniDocBench at 95.38(3 posts)→
More from Models
- System One Models Like Jev as Primitives for Frontier Agents — omarsar0 · 2026-09-20
- Code-only heuristic policies can beat frontier models on Craftax, evals researcher says — JoshPurtell · 2026-09-20
- Codex usage reset now live for all, big OpenAI release teased for Tuesday — kimmonismus · 2026-09-20
- Jev beats GPT-5.6 Luna on PR review: 1.93x faster at $0.0014 per run — aniketmaurya · 2026-09-20
- Fruit fly connectome chess model beats Jev 4-1 in 10 games, with a playable demo site — maximelabonne · 2026-09-20
- Bonsai 2 27B safety guardrails reportedly cut SWE-bench and Terminal-bench scores by ~20 points — julianharris · 2026-09-20