Tencent Hunyuan maps scaling laws for native multimodal pretraining from scratch
Tencent-Hunyuan · hf · 2026-07-27
- Tencent Hunyuan studies how native multimodal pre-training from scratch scales under a fixed compute budget.
- The paper finds that minimal loss follows a predictable compute law, while compute-optimal model size and token count follow power laws.
- Language and multimodal objectives behave differently: the language allocation law is mostly invariant to data composition, but the multimodal allocation law is highly sensitive to it.
- Text-heavy mixtures only become compute-efficient at larger model scales, pushing the optimal allocation toward more model capacity.
- The authors also derive an efficiency frontier that specifies model size, token count, and data mixture, and report positive cross-modal transfer to pure-text spatial reasoning and multimodal in-context learning.
More from Multimodal
- Skywork Video says AI video must move from prompting to full production workflows — Shruti_0810 · 2026-07-27
- Skywork Video adds brand guidelines so logos, fonts, and colors can be reused — Shruti_0810 · 2026-07-27
- TapNow’s CreativeOS powers a Shenzhen AI horror hackathon and pushes video creation into one workflow — 葬AI · 2026-07-27
- A fine-tuned Krea 2 raw model produced a rainy-night driving scene — darlens13 · 2026-07-27
- Users ask whether Video2X can load custom OpenModelDB models — Used-Profit2355 · 2026-07-27
- A builder wants AI to reverse-engineer viral video effects into ComfyUI workflows — stale2000 · 2026-07-27