1B Vision Model Beats Google and NVIDIA in Depth Estimation

thetripathi58 · x · 2026-07-08

According to the post, the 1B vision model LingBot-Vision, trained by Robbyantbrain, outperforms Google SigLIP 2 and NVIDIA AM-RADIO on the NYU-Depth v2 depth estimation benchmark with a lower RMSE.

The author emphasizes that this performance gap primarily stems from the pre-training objectives rather than just model scale.

Related event: Ant Robbyant Open-Sources LingBot Vision Models, Topping Depth Benchmarks(18 posts)→

Original post →

More from Embodied

Embodied channel →