Tencent Hunyuan Maps Scaling Laws for Encoder-Free Multimodal Pretraining
Tencent-Hunyuan · hf · 2026-09-29
Tencent Hunyuan systematically compares scaling laws for encoder-free vs encoder-based multimodal LLMs. Key findings: removing the visual encoder shifts compute-optimal allocation toward larger models for the multimodal objective; encoder-free lags at small scale but is predicted to catch up around 10^22 FLOPs, within practical pretraining budgets; and the LM learns to take over the encoder's role, with bidirectional visual-token interactions, earlier-layer visual processing, and more concentrated expert routing as compute grows. Encoder-free architectures emerge as a promising direction.
More from Models
- AI's verbosity problem: 'simple is always best' as users tire of walls of text — djcows · 2026-09-29
- Users say Claude's $200 plan was quietly halved, framed as a long-term improvement — gandamu_ml · 2026-09-29
- 'Opus 5 was a demon': users say Opus 5.5 is markedly better behaved — gandamu_ml · 2026-09-29
- Dev switches from Cursor+Grok 4.7 to Claude Code+Opus 5.5 for a full day — jonathan_wilke · 2026-09-29
- Same prompt showdown: Opus 5.5 Max effort vs Grok 4.7 Fast xhigh for motion graphics video — FinanceYF5 · 2026-09-29
- Asking Opus "what do you think?" gets flagged as a distillation attack — Dany0 · 2026-09-29