Meta/Oxford study: multimodal models need only 5% image-generation data, language training is key
rohanpaul_ai · x · 2026-08-14
A study by Meta and Oxford University finds that multimodal models may need surprisingly little image-generation data if language and visual understanding are trained together from the start. Language training already helps with vision, and learning to understand images also improves generation, but training to generate images does little for language or understanding. In 1T-token experiments, the best split was 70% language, 25% image understanding, and only 5% image generation. At 13.5B scale over 2T tokens, even with 5x fewer image-generation tokens, GenEval improved from 0.467 to 0.482, with language and understanding also improving. The study warns against adding vision too late, as models may develop 'vision laziness' and rely on language shortcuts. Paper: 'Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes'.
More from Multimodal
- Running Minimax H3 on DGX Spark: 5-Second Video Costs Only $0.001 — GabryIta · 2026-08-14
- Newbie asks: Is SDXL still worth learning? — Foreign-Roof4913 · 2026-08-14
- Minimax H3 Ref2VA Lipsync workflow released — Most_Way_9754 · 2026-08-14
- MiniMax Music 3 Goes Open Source, Community Predicts Music Generation Revolution — TheChuckTone · 2026-08-14
- MiniMax H3 Text-to-Video Test: Bowling Scene Looks Impressive — howdyquade · 2026-08-14
- Text-to-Music Models Fail at Hardstyle Kick: A New Benchmark — cephaloform · 2026-08-14