Ant Group Open-Sources LingBot Vision: 1.1B Params Beat 7B DINOv3

AdinaYakup · x · 2026-07-07

Ant Group has released the LingBot Vision series of self-supervised visual backbone networks under the Apache 2.0 license, with model scales ranging from ViT-S to ViT-g. The core innovation is the "masked boundary modeling" pre-training approach, which enables the model to maintain sharper feature representations at object edges, making it highly suitable for dense spatial perception tasks. Self-reported evaluations show that the 1.1B parameter version outperforms the 7B parameter DINOv3 on the NYU-Depth v2 depth estimation task, demonstrating a significant leap in parameter efficiency.

Related event: Ant Robbyant Open-Sources LingBot Vision Models, Topping Depth Benchmarks(18 posts)→

Original post →

More from Research

Research channel →