NTU and SenseTime Propose Native Unified Vision Model NEO-ov

jiqizhixin · x · 2026-07-05

NTU and SenseTime Research proposed NEO-ov, a native unified vision model that learns pixel-to-text correspondences end-to-end without requiring external encoders or fusion tricks. The model rivals modular architectures on standard tasks and excels in fine-grained visual perception across single-image, multi-image, and video inputs, proving that native architectures remain competitive at scale. Both the paper and code have been released.

Original post →

More from Multimodal

Multimodal channel →