Tencent Hunyuan Maps Scaling Law for Encoder-Free Multimodal Models
Tencent's Hunyuan team systematically compared encoder-free and encoder-based multimodal LLMs, finding that encoder-free models overtake their counterparts beyond 10^22 FLOPs of pretraining compute, suggesting vision encoders can be dropped at sufficient scale.
2026-09-29 ~ 2026-09-30 · 2 related posts
- Tencent Hunyuan Maps Scaling Laws for Encoder-Free Multimodal Pretraining — Tencent-Hunyuan · 2026-09-29
- You can drop the vision encoder once pretraining compute exceeds 1e22 FLOPs — heghbalz · 2026-09-30