Tencent open-sources Youtu-Parsing-Omni, a 5B omni-modal parsing model

jacek2023 · reddit · 2026-10-09

Tencent released Youtu-Parsing-Omni, a compact 5B model that turns a single input — document page, natural image, chart/flowchart, geometry figure, audio, or audio-visual video — into one structured JSON covering both perception (layout, OCR/ASR, tables, formulas, bboxes) and cognition (captions, reports), switched by task prompt. It hits 96.96 on OmniDocBench v1.6 (SOTA), 75.08 on OmniParsingBench (best open-weight, behind only Gemini-3-Pro), and ships with a vLLM plugin and pinned serving configs.

Related event: Tencent Open-Sources 5B Omni-Modal Parsing Model(2 posts)→

Original post →

More from Models

Models channel →