Embodied Vision Model Leads in Depth and Segmentation

thetripathi58 · x · 2026-07-08

The post claims that while SigLIP 2 performs better on ImageNet classification, dense spatial tasks are more critical for robots.

LingBot-Vision improves depth estimation by about 40% compared to Google's model, while also gaining 10.8 mIoU on the ADE20K segmentation benchmark.

Related event: Ant Robbyant Open-Sources LingBot Vision Models, Topping Depth Benchmarks(18 posts)→

Original post →

More from Embodied

Embodied channel →