TDDN Fuses DINOv3 with CleanDIFT, More Than Triples CLIP's Dense Prediction Accuracy
udmrzn · x · 2026-09-11
arXiv paper 2609.07937: CLIP-ViT-based VLMs trade fine-grained perception for semantics, hurting structured visual reasoning like puzzles. TDDN fuses DINOv3's global semantics with CleanDIFT's fine-grained spatial features into a DiffusedDINO encoder aligned to RoBERTa-L.
Results with frozen backbones and only 590K alignment pairs:
- Matches CLIP on image-text retrieval, beating it in three of four settings;
- More than triples CLIP's dense prediction accuracy (ADE20K 5.20→18.11 mIoU, COCO-Stuff 7.35→24.44);
- Leads segmentation benchmarks among general-purpose contrastive encoders including SigLIP 2;
- Doubles CLIP's segmentation accuracy (11.04→22.51 mIoU) on the newly introduced Puzzle Perception dataset.
More from Multimodal
- Curated collection of 153 GPT Astra prompts showcases one-shot 3D and game generation — TheMoonMidas · 2026-09-11
- DeepSeek V4.1 Flash multimodal: 45T image-text tokens and modality-level load balancing — nrehiew_ · 2026-09-11
- Is local AI video upscaling still broken? Hailuo 768p to 2K/4K求助 — Infinite-Emptiness · 2026-09-11
- Astra Can Compose SNES-Style Chiptunes, Vibe Coders Report — AIandDesign · 2026-09-11
- H3 Max video endpoint teaser blurs the line between real and AI footage — noahsolomon · 2026-09-11
- 14-Year Editor Tests GPT-6 Astra: It Watched 84 Clips, Cut, Scored and Subtitled in DaVinci Resolve — TheMoonMidas · 2026-09-11