NEO原生多模态架构:从零训练挑战视觉编码器 | 论文发布
liuziwei7 · x · 2026-08-16
A new paper titled "From Pixels to Words" proposes principles for constructing native Vision-Language Models (VLMs) and introduces the NEO model family. Addressing the controversy over VLM vision encoders, the authors argue that the best encoder is no encoder and that training from scratch is always superior. NEO models align pixel and word representations within a shared semantic space, integrating separate vision and language strengths. Trained from scratch on 390M image-text pairs, NEO significantly narrows the gap with top-tier modular counterparts.
More from Multimodal
- First animation created with MiniMax H3 Ref2VA in ComfyUI — NobodySnJake · 2026-08-16
- MiniMax H3 acts as multi-ref image editor with a Tamagotchi judge — RobbaW · 2026-08-16
- AI-generated 'The Legend of Wu Zetian' music video showcases 36 shots — Hongyi_AI · 2026-08-16
- Creating a 6-Minute Animation with MiniMax H3: Workflow & Benchmarks — dassiyu · 2026-08-16
- Fan-made video with Minimax H3: Clawhauser asks Sonic about Amy in Zootopia style — TigerClaw305 · 2026-08-16
- Seedance 2.5 launches on TapNow with native 1080p video generation — SimplyAnnisa · 2026-08-16