You can drop the vision encoder once pretraining compute exceeds 1e22 FLOPs
heghbalz · x · 2026-09-30
rosinality highlights a research finding: once multimodal pretraining compute exceeds 1e22 FLOPs, you can drop the separate vision encoder entirely — suggesting that at sufficient scale, native multimodal training absorbs what the vision encoder normally provides.
Related event: Tencent Hunyuan Maps Scaling Law for Encoder-Free Multimodal Models(2 posts)→
More from Models
- DeepSeek's forgotten male persona: why the 'big fat fish' meme won the Bilibili war — teortaxesTex · 2026-09-30
- GPT-6.1 Sol appears silently nerfed mid-task, user reports 3B-level output quality — Ferzelibey · 2026-09-30
- mradermacher quants get Gemma 26B to 75 tok/s on 2x RTX 4060 8GB — Spiritual_Impress_30 · 2026-09-30
- DeepSeek is giving users 6 yuan in free API credits via its harness — teortaxesTex · 2026-09-30
- User claims OpenAI bots autonomously scan your Gmail after connecting and keep the data — alexcovo_eth · 2026-09-30
- Rumor: DeepSeek's rumored single-GPU model may have been trained on Ascend — teortaxesTex · 2026-09-30