ColNanoVDR distills multi-vector visual document retrieval without documents, keeping 95% NDCG@5 at 149M params

nanovdr · hf · 2026-09-29

Multi-vector visual document retrieval (VDR) retrievers lead the field but run a multi-billion-parameter query encoder on every search, and standard score distillation requires caching terabytes of page tokens.

ColNanoVDR claims to be the first framework to bring document-free distillation to multi-vector VDR: the student trains purely on the teacher's query embeddings, never touching pages. Its OTW objective (Optimal Transport with Learned Weights) aligns student query tokens to the teacher via entropic optimal transport with a learned per-token weight, requiring no token correspondence, and the authors prove this alignment cost bounds the MaxSim score difference on every page.

Results: 149M text-only students distilled from five SOTA teachers retain 95% of teachers' NDCG@5 on ViDoRe v1-v3 while encoding queries up to 26x faster; under identical training OTW matches score distillation with zero page encoding and 12.6x less cached teacher data read.

Related event: ColNanoVDR Distills Multi-Vector Doc Retrieval Without Encoding Documents(2 posts)→

Original post →

More from Research

Research channel →