Ant Group Releases LingBot-Vision Self-Supervised Vision Backbone
Simple_Response8041 · reddit · 2026-07-07
Ant Group has released LingBot-Vision, a DINO-family self-supervised vision backbone available in four sizes (ViT-S/B/L/g, 21M to 1.1B) under the Apache-2.0 license. The core innovation is boundary-driven masking: a teacher network predicts object boundaries and forces these tokens into the student network's mask, preventing it from simply copying flat contexts for reconstruction. The entire process requires no labels, text supervision, or external edge detectors. On NYUv2 depth estimation, the 1.1B flagship achieves an optimal 0.296 RMSE (compared to 0.309 for DINOv3-7B), while the 0.3B ViT-L reaches 0.310, rivaling the 7B model with roughly 23x fewer parameters. Its weak point is ImageNet classification, where both the flagship and L sizes trail DINOv3. The repository includes a lightweight loader to output frozen backbones for feature extraction, focusing on dense features like depth, segmentation, and tracking rather than conversational tasks.
Related event: Ant Robbyant Open-Sources LingBot Vision Models, Topping Depth Benchmarks(18 posts)→
More from Research
- Nat Lambert shares a reading list on synthetic data and agentic SFT data — natolambert · 2026-07-22
- Lightwheel AI Launches SimReadyGen: Text-to-Physics-Accurate Robot Sim Assets — ZeYanjie · 2026-07-22
- PNAS special issue examines copyright, governance, and AI in the legal system — chrmanning · 2026-07-22
- WeirdChat catalogs strange model behaviors from more than 100 million sampled responses — JacobSteinhardt · 2026-07-22
- New agentic benchmark shows AI managers escalate to coercion and fake success — Jasmine Brazilek · 2026-07-22
- Ai2’s Asta adds one-click handoff and self-checking deep paper search — allen_ai · 2026-07-22