Jina AI's jina-ocr-v1 document understanding model trends on Hugging Face
jinaai · hf · 2026-09-19
jina-ocr-v1 from Jina AI is trending on Hugging Face. The image-text-to-text multimodal model (built on deepseekvlv2) targets OCR, multilingual document intelligence, and vision-language understanding, shipped with transformers and safetensors support for document parsing workflows.
More from Multimodal
- Alibaba's Qwen3.8-LiveTranslate does real-time speech translation with 2.3s lag and speaker separation — xiaohu · 2026-09-19
- Alibaba's Qwen3.8-LiveTranslate cuts speech translation lag to 2.3s, adds speaker separation — xiaohu · 2026-09-19
- Fully open-source Diffusion Studio generated a 450K-view launch video without touching the timeline — _AustinCalvert_ · 2026-09-19
- Over half of AI videos will be rendered from code, predicts Diffusion Studio demo — _AustinCalvert_ · 2026-09-19
- Early tests of GPT image model show stunning 70s sci-fi style 9-grid portraits — aziz4ai · 2026-09-19
- Codex ships with Images 2.5 built in: redesign pages with imagegen first, then implement — OpenAIDevs · 2026-09-19