Ant Lingbo Open-Sources LingBot-Vision, a Boundary-Centric Spatial Vision Model

机器之心 · wechat · 2026-07-07

Ant Lingbo has released and open-sourced the visual foundation model LingBot-Vision (approx. 1.1B parameters) and the depth model LingBot-Depth2.0, along with the technical report "Vision Pretraining for Dense Spatial Perception" and the code. The core method, "boundary-centric masked modeling," involves real-time prediction of object boundaries during training and forced masking, compelling the model to reconstruct geometric structures using context. It filters pseudo-boundaries using a-contrario testing and solves the cold start problem with bootstrapped corners. This approach requires only about 1/10 of the data and less than 1/3 of the training compute.

Related event: Ant Robbyant Open-Sources LingBot Vision Models, Topping Depth Benchmarks(18 posts)→

Original post →

More from Embodied

Embodied channel →