Tencent's WeVisDoc tops OmniDocBench with 95.38 via two-stage data-centric training
tencent · hf · 2026-09-18
Tencent introduces WeVisDoc, a two-stage data-centric framework for robust end-to-end document parsing. Stage I broadens semantic, structural, and appearance coverage via heterogeneous data and structure-preserving degradation synthesis; Stage II uses a held-out probe to diagnose residual errors within visual-structural clusters and reallocates the target-token budget accordingly. WeVisDoc-4B scores 95.38 overall on OmniDocBench v1.6 and 75.54 mean across three PureDocBench tracks, ranking first in all four settings, with the largest gains on degraded inputs (e.g., +4.03 on Real Degraded for the 4B model).
More from Research
- ML is leaving its alchemy era: assumed truths can now be instrumented and falsified — mike64_t · 2026-09-18
- Together AI paper: score centering cancels drift in off-policy RL under train-inference mismatch — PandaAshwinee · 2026-09-18
- Cell's AI-in-biology special issue debuts LongevityBench, a 17-task aging benchmark — JosephJacks_ · 2026-09-18
- After OpenAI cracks Navier-Stokes, the era of the human mathematician is ending — yeastsplainer · 2026-09-18
- Custom BB-SLAM hits 0.02m accuracy with pure visual odometry, no LiDAR or IMU — broodsugar · 2026-09-18
- AI4Science Breakthrough Said to Exceed Expectations, Could Transform Computational Chemistry — BenBlaiszik · 2026-09-18