Tencent releases Youtu-Parsing-Omni, an open multimodal model for document parsing, OCR, audio and video

tencent · hf · 2026-10-09

Tencent has released Youtu-Parsing-Omni on Hugging Face, an image-text-to-text multimodal model built on the youtuvita architecture.

The model targets universal document parsing, combining OCR with audio and video understanding. Weights are available in transformers/safetensors format for direct loading via the transformers library, making it useful for building document understanding and multimodal extraction pipelines.

Original post →

More from Models

Models channel →