Tencent open-sources Youtu-Parsing-Omni, a 5B omni-modal parsing model
jacek2023 · reddit · 2026-10-09
Tencent released Youtu-Parsing-Omni, a compact 5B model that turns a single input — document page, natural image, chart/flowchart, geometry figure, audio, or audio-visual video — into one structured JSON covering both perception (layout, OCR/ASR, tables, formulas, bboxes) and cognition (captions, reports), switched by task prompt. It hits 96.96 on OmniDocBench v1.6 (SOTA), 75.08 on OmniParsingBench (best open-weight, behind only Gemini-3-Pro), and ships with a vLLM plugin and pinned serving configs.
Related event: Tencent Open-Sources 5B Omni-Modal Parsing Model(2 posts)→
More from Models
- Elliot Glazer: Astra flags flawed proof in OpenAI's Weil classes paper — burny_tech · 2026-10-09
- Specialist raises concerns over OpenAI's Birch–Swinnerton-Dyer related results — burny_tech · 2026-10-09
- Mathematicians' group AHM slams OpenAI's 700-file release: 'power, not scholarship' — ChrSzegedy · 2026-10-09
- Mathematicians find issues beyond sloppy presentation in OpenAI's math dump — burny_tech · 2026-10-09
- OpenAI releases 372 math proof claims, including 23 Erdős problems spanning 1268 pages — burny_tech · 2026-10-09
- dhh prefers Codex as main coding model with Claude secondary, praises Sol 6.1 — npew · 2026-10-09