DINOv2 and Qwen3 embedding spaces aligned with zero image-caption pairs
TimDarcet · x · 2026-10-10
Dominik Schnaus demonstrates aligning DINOv2 and Qwen3 embedding spaces without a single image-caption pair — DINOv2 has never seen captions and Qwen3 never images — and it works even when images and captions come from different datasets. Project page and code available.
Related event: DINOv2 and Qwen3 embeddings aligned without any image-text pairs(4 posts)→
More from Research
- SF event Oct 22 asks what it takes for AI to run scientific research end to end — nlarusstone · 2026-10-10
- Paper argues nested von Neumann architecture can make a million processors act as one computer — bronzeagepapi · 2026-10-10
- UPenn paper unifies diffusion and autoregression on one corruption lattice to predict decoding costs — upenn · 2026-10-10
- CMU's Keenan Crane wraps AI-assisted 3D modeling thread, releases all model files under CC0 — keenanisalive · 2026-10-10
- From Maxwell's field lines to Jim Blinn's blobby molecular visualization — keenanisalive · 2026-10-10
- OmniSeek turns Omni-LLMs into evidence-seeking agents, +15.5 points on VideoHolmes — mohitban47 · 2026-10-10