Towards the Physics of Multimodal Pretraining: Insights from a 13.5B MoE Model
khademinori · x · 2026-08-07
While most LLMs start with language pretraining, native multimodal training is often avoided due to instability and modality competition. This research systematically unpacks the underlying mechanics of multimodal pretraining across four key aspects:
- Knowledge Flow: Cross-modal transfer across language, understanding, and generation is highly asymmetric and concept-dependent.
- Modality Synergy: Task complexity and architectural sharing dictate whether modalities help or compete with each other.
- Early Unification: Early, simultaneous unification outperforms late or sequential training. Delayed visual integration can induce "vision laziness" in models.
- Recipes & Scale: The insights from controlled studies are used to build efficient pretraining recipes, which are validated at scale by training a 13.5B Mixture-of-Experts (MoE) model on 2T tokens.
Related event: Meta Unveils Mechanisms of Multimodal Pre-training(4 posts)→
More from Research
- Netflix details its production LLM judge: hundreds of thousands of recommendations scored weekly — omarsar0 · 2026-08-24
- Nature Comment: Provenance, not interpretability, grounds trust in autonomous science — gabepgomes · 2026-08-24
- New Architecture RHEA: Train 1B Model on 8GB VRAM — zemondza · 2026-08-24
- Trained two 16M-param models to do generative CAD with real physics — debreuil · 2026-08-24
- Claude model helps discover complex structure on S^6, solving 60-year-old math problem — Singularitarian · 2026-08-24
- Study: Agents read instructions/notes 60.5% of the time, rarely touch API docs — dair_ai · 2026-08-24