1B Vision Model Beats 7B in Depth Tasks

rohanpaul_ai · x · 2026-07-09

Robbyant released LingBot-Vision, claiming it as the first "spatial-native" vision foundation model, shifting its training focus from semantic recognition to boundaries and spatial relationships. The author emphasizes that this is crucial for robotics, as robots need to perceive spatial details like object boundaries, transparent items, and thin cables, rather than just recognizing a "cat" or a "chair."

The post also shares benchmark results: this 1B parameter model achieved an RMSE of 0.296 on NYU-Depth v2, outperforming the 7B model DINOv3-7B (0.309)—and it was achieved with a frozen model, single linear head, and zero fine-tuning.

Related event: Robbyant Releases LingBot-Vision: 1B Spatial Model Beats 7B(7 posts)→

Original post →

More from Embodied

Embodied channel →