Ant's Robbyant Releases LingBot-Depth 2.0 Research
Ok-Line2658 · reddit · 2026-07-07
Robbyant, an embodied AI company under Ant Group, proposes LingBot-Depth 2.0. The core idea is to treat the sensor's own invalid depth regions (specular highlights, transparent objects, textureless areas) as masking signals, rather than using random patch masking. This forces the model to learn directly from the real failure distributions it will encounter during inference.
Version 2.0 only changes the encoder initialization and data scale, keeping the rest of the training recipe unchanged. Controlled experiments show that LingBot-Vision initialization outperforms on almost all ViT-L benchmarks and most ViT-g benchmarks, with the advantage widening as data scales up. The report achieves the best RMSE in 7 out of 8 patch masking/sparse benchmarks, performing strongest on the ClearGrasp transparent object dataset, and roughly halving the DIODE indoor patch masking RMSE compared to version 1.0.
Related event: Ant Robbyant Open-Sources LingBot Vision Models, Topping Depth Benchmarks(18 posts)→
More from Embodied
- Amazon and Google sold 600M+ smart speakers, so why no AGI-era successor? — julianlehr · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- Polish developers build iPhone app that detects nearby Meta smart glasses — Low-Honeydew6483 · 2026-09-11
- Ant's Afu health AI hits 150M users, unveils AI+hardware health alliance at Bund Summit — APPSO · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11