TDDN Fuses DINOv3 with CleanDIFT, More Than Triples CLIP's Dense Prediction Accuracy

udmrzn · x · 2026-09-11

arXiv paper 2609.07937: CLIP-ViT-based VLMs trade fine-grained perception for semantics, hurting structured visual reasoning like puzzles. TDDN fuses DINOv3's global semantics with CleanDIFT's fine-grained spatial features into a DiffusedDINO encoder aligned to RoBERTa-L.

Results with frozen backbones and only 590K alignment pairs:

Original post →

More from Multimodal

Multimodal channel →