Tencent Hunyuan Maps Scaling Laws for Encoder-Free Multimodal Pretraining

Tencent-Hunyuan · hf · 2026-09-29

Tencent Hunyuan systematically compares scaling laws for encoder-free vs encoder-based multimodal LLMs. Key findings: removing the visual encoder shifts compute-optimal allocation toward larger models for the multimodal objective; encoder-free lags at small scale but is predicted to catch up around 10^22 FLOPs, within practical pretraining budgets; and the LM learns to take over the encoder's role, with bidirectional visual-token interactions, earlier-layer visual processing, and more concentrated expert routing as compute grows. Encoder-free architectures emerge as a promising direction.

Original post →

More from Models

Models channel →