Tencent releases Youtu-Parsing-Omni, an open multimodal model for document parsing, OCR, audio and video
tencent · hf · 2026-10-09
Tencent has released Youtu-Parsing-Omni on Hugging Face, an image-text-to-text multimodal model built on the youtuvita architecture.
The model targets universal document parsing, combining OCR with audio and video understanding. Weights are available in transformers/safetensors format for direct loading via the transformers library, making it useful for building document understanding and multimodal extraction pipelines.
More from Models
- Developer says he now manages 90% of his coding work through Grok bot — jonathan_wilke · 2026-10-09
- State of AI Report 2026 Predicts AGI by 2027 and Lists 9 Bets for the Next Year — FinanceYF5 · 2026-10-09
- Anthropic Leads on Coding but Not Overall: Astra Still Ahead, Astra 6.1 Coming — haider1 · 2026-10-09
- ChatGPT's $500 tier measured: weekly cap hits at ~90M Sol + 25M Luna tokens — airesearch12 · 2026-10-09
- Opus Now Imitates Your Voice in Video Easier Than Writing Human-Sounding Prose — tobowers · 2026-10-09
- Jev works surprisingly well as a cheap general-purpose classifier, dev reports — max_paperclips · 2026-10-09