Ant Lingbo Open-Sources LingBot-Vision, a Boundary-Centric Spatial Vision Model
机器之心 · wechat · 2026-07-07
Ant Lingbo has released and open-sourced the visual foundation model LingBot-Vision (approx. 1.1B parameters) and the depth model LingBot-Depth2.0, along with the technical report "Vision Pretraining for Dense Spatial Perception" and the code. The core method, "boundary-centric masked modeling," involves real-time prediction of object boundaries during training and forced masking, compelling the model to reconstruct geometric structures using context. It filters pseudo-boundaries using a-contrario testing and solves the cold start problem with bootstrapped corners. This approach requires only about 1/10 of the data and less than 1/3 of the training compute.
Related event: Ant Robbyant Open-Sources LingBot Vision Models, Topping Depth Benchmarks(18 posts)→
More from Embodied
- Amazon and Google sold 600M+ smart speakers, so why no AGI-era successor? — julianlehr · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- Polish developers build iPhone app that detects nearby Meta smart glasses — Low-Honeydew6483 · 2026-09-11
- Ant's Afu health AI hits 150M users, unveils AI+hardware health alliance at Bund Summit — APPSO · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11