LingBot-Vision: Self-supervised ViT backbones via masked boundary modeling
tom_doerr · x · 2026-08-23
LingBot-Vision introduces a family of self-supervised Vision Transformer backbones for dense spatial perception, ranging from ViT-S/16 to a 1.1B-parameter ViT-g/16.
Key Features:
- Pretrained using Masked Boundary Modeling.
- The objective encourages spatially structured patch features while retaining strong semantic representations.
- The model simultaneously learns boundaries, shapes, and semantic regions.
The repository includes code, examples, and an accompanying paper.
More from Research
- Stanford/PKU paper: never train inside the world model — QWM uses it only for action selection — rohanpaul_ai · 2026-08-23
- 4B Model BFCL Jumps 15%: Snowflake Proves Power of Mid-Tool Training — TheTuringPost · 2026-08-23
- MemFail Paper: Stress-Testing Failure Modes of LLM Memory Systems — xuandongzhao · 2026-08-23
- Study: Why one AI is often better than four for decision making — Exponential View (Azeem Azhar) · 2026-08-23
- Revisiting Classics: Progress is Cumulative, Not Revolutionary — DJiafei · 2026-08-23
- NeurIPS 2026 to host BabyVLM workshop on learning like babies — LChoshen · 2026-08-23