Ant LingBot-Depth2.0 Tops 12 World Rankings, Open-Sources Embodied AI Vision Base

新智元 · wechat · 2026-07-07

Ant LingBot released the LingBot-Depth2.0 spatial perception model, securing 12 global first places across public and private datasets. It tackles core embodied AI challenges like depth estimation for glass, mirrors, and transparent objects, alongside fine object detection and long-range perception. Abandoning general vision models like DINOv3, the team built LingBot-Vision from scratch—the world's first "spatial-native" vision foundation model for embodied AI. With just 1.1 billion parameters, it outperforms 7B-level competitors. LingBot-Vision is open-sourced on Hugging Face, ModelScope, and GitHub, with its technical report on arxiv. Commercially, Ant LingBot partnered with Orbbec to launch the EGO-RGBD data collection device with integrated SDK, completing the loop from model capability to commercial deployment.

Related event: Ant Robbyant Open-Sources LingBot Vision Models, Topping Depth Benchmarks(18 posts)→

Original post →

More from Embodied

Embodied channel →