Aligning DINOv2 and Qwen3 embedding spaces with zero image-caption pairs
AccBalanced · x · 2026-10-11
dominikschnaus presents research aligning the embedding spaces of DINOv2 (which has never seen a caption) and Qwen3 (which has never seen an image) without a single image-caption pair — it even works when images and captions come from different datasets. Project page and paper are available.
More from Multimodal
- Extending Qwen Image 2.1 Turbo's magic 8-step sigmas to 12-14 steps — GTManiK · 2026-10-11
- Seedance 2 Video Demo: The Missing Piece for Your Computer — Ok-Nerve941 · 2026-10-11
- ComfyUI on a $899 M4 Mac Mini 16GB: Full Benchmark Results and Workflows — FaatmanSlim · 2026-10-11
- Prime Video picks up AI-generated spin-off of Turkish hit Çukur, every frame AI-made — lmoroney · 2026-10-11
- Blogger warns: most AI-generated product demo videos are deceptive — lipeng0820 · 2026-10-11
- Grok Imagine hands-on: transitions still hit-or-miss but a year of progress shows — RachelVT42 · 2026-10-11