Meta/Oxford study: multimodal models need only 5% image-generation data, language training is key

rohanpaul_ai · x · 2026-08-14

A study by Meta and Oxford University finds that multimodal models may need surprisingly little image-generation data if language and visual understanding are trained together from the start. Language training already helps with vision, and learning to understand images also improves generation, but training to generate images does little for language or understanding. In 1T-token experiments, the best split was 70% language, 25% image understanding, and only 5% image generation. At 13.5B scale over 2T tokens, even with 5x fewer image-generation tokens, GenEval improved from 0.467 to 0.482, with language and understanding also improving. The study warns against adding vision too late, as models may develop 'vision laziness' and rely on language shortcuts. Paper: 'Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes'.

Original post →

More from Multimodal

Multimodal channel →