You can drop the vision encoder once pretraining compute exceeds 1e22 FLOPs

heghbalz · x · 2026-09-30

rosinality highlights a research finding: once multimodal pretraining compute exceeds 1e22 FLOPs, you can drop the separate vision encoder entirely — suggesting that at sufficient scale, native multimodal training absorbs what the vision encoder normally provides.

Related event: Tencent Hunyuan Maps Scaling Law for Encoder-Free Multimodal Models(2 posts)→

Original post →

More from Models

Models channel →