NEO原生多模态架构:从零训练挑战视觉编码器 | 论文发布

liuziwei7 · x · 2026-08-16

A new paper titled "From Pixels to Words" proposes principles for constructing native Vision-Language Models (VLMs) and introduces the NEO model family. Addressing the controversy over VLM vision encoders, the authors argue that the best encoder is no encoder and that training from scratch is always superior. NEO models align pixel and word representations within a shared semantic space, integrating separate vision and language strengths. Trained from scratch on 390M image-text pairs, NEO significantly narrows the gap with top-tier modular counterparts.

Original post →

More from Multimodal

Multimodal channel →