NTU and SenseTime Propose Native Unified Vision Model NEO-ov
jiqizhixin · x · 2026-07-05
NTU and SenseTime Research proposed NEO-ov, a native unified vision model that learns pixel-to-text correspondences end-to-end without requiring external encoders or fusion tricks. The model rivals modular architectures on standard tasks and excels in fine-grained visual perception across single-image, multi-image, and video inputs, proving that native architectures remain competitive at scale. Both the paper and code have been released.
More from Multimodal
- Creator turns Bahamut vs Tiamat rivalry into an AI cinematic battle with Midjourney, GPT Image 2 and Seedance — azed_ai · 2026-09-11
- invideo launches AI agent-powered editor to automate repetitive editing tasks — azed_ai · 2026-09-11
- fable 5.1 recreates The Starry Night with 256,157 JavaScript brush strokes — cedric_chee · 2026-09-11
- GPT-6 Astra + Hyper3D Rodin MCP Generates 3D Assets in One Agent Flow — ahuja_priyank · 2026-09-11
- Five tells that still make AI video read as AI, from physics glitches to missing operators — NewPhoneWhotiz · 2026-09-11
- MiniMax + ComfyUI used to create Michael Jackson moonwalk-on-the-moon short film — Inside-Cantaloupe233 · 2026-09-11