Tencent Hunyuan Open-Sources 1B End-to-End OCR
腾讯混元 · wechat · 2026-07-13
Tencent Hunyuan released HyOCR-1.5, a 1B-parameter end-to-end OCR expert model with a fully open-source stack covering training, inference, and weights. It supports over 8 text-centric tasks and leverages DFlash-based speculative decoding to achieve up to a 6.37× inference speedup. It also scores 94.74 on OmniDocBench v1.6, ranking among the top tier of end-to-end OCR models.
The article highlights three key aspects:
- Faster: Uses a lightweight draft model with verification-based speculative decoding to accelerate autoregressive generation for long documents, delivering significant gains for long outputs and table pages.
- Stronger: Introduces AgenticDataFlow, enabling agents to automatically collect materials, develop data pipelines, and iterate based on model weaknesses, thereby improving capabilities in low-resource languages, ancient texts, and multi-image QA.
- More Comprehensive: Integrates 4K resolution, 128K context, and RL on the training side; evaluations cover document parsing, ancient texts, charts, multilingual scenarios, multi-image QA, and hallucusion suppression.
The article also shares multiple results: achieving SOTA in ancient text recognition at the 1B scale on Chronicles-OCR; performing close to or even exceeding some 8B models on ChartArena; and reaching a 99.8% accuracy in text-free image processing.
More from Multimodal
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22
- Reddit user seeks ComfyUI NSFW text-to-image and image-to-video workflows under 20 GB VRAM — hobbyist2020 · 2026-07-22
- Krea 2 users recommend a two-pass Clownshark sampler setup for sharper image details — listopalafoto · 2026-07-22
- Gemini Omni Flash turns a boat cabin into a cave in Flow by Google — chrisfirst · 2026-07-22
- A simple workflow to turn a photo into an image prompt using Gemini, Grok, or GPT Image — harshitagu72595 · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22