LingBot-Vision Boundary Mask Self-Supervised Pretraining

StillThese3747 · reddit · 2026-07-07

This paper proposes an online boundary field-driven masked self-supervised pretraining method. A teacher model predicts dense boundary fields online, and tokens located on these boundaries are forcibly included in the student mask, compelling the student to reconstruct regions that cannot be inferred from context. The boundary targets are generated by the teacher itself, rather than relying on labels or external edge detectors. In its design, the boundary field is recast as a pixel-wise classification distribution to reuse the centering/sharpening mechanisms of self-distillation, preventing representation collapse. Additionally, the decoded segmentation must pass an a-contrario validation before participating in the supervision.

Self-reported data (all reported by the authors): A 1.1B (patch-16) model achieves an RMSE of 0.296 on the NYUv2 linear probe, outperforming DINOv3-7B's 0.309 in their comparisons. A distilled ViT-L (0.3B) reaches 0.310, approaching the 7B level. The data budget uses only 161 million images, less than a third of DINOv3's. Weaknesses: It lags behind DINOv3 in giant/L scale ImageNet classification and on ADE20K.

Related event: Ant Robbyant Open-Sources LingBot Vision Models, Topping Depth Benchmarks(18 posts)→

Original post →

More from Research

Research channel →