ColNanoVDR distills multi-vector visual document retrieval without documents, keeping 95% NDCG@5 at 149M params
nanovdr · hf · 2026-09-29
Multi-vector visual document retrieval (VDR) retrievers lead the field but run a multi-billion-parameter query encoder on every search, and standard score distillation requires caching terabytes of page tokens.
ColNanoVDR claims to be the first framework to bring document-free distillation to multi-vector VDR: the student trains purely on the teacher's query embeddings, never touching pages. Its OTW objective (Optimal Transport with Learned Weights) aligns student query tokens to the teacher via entropic optimal transport with a learned per-token weight, requiring no token correspondence, and the authors prove this alignment cost bounds the MaxSim score difference on every page.
Results: 149M text-only students distilled from five SOTA teachers retain 95% of teachers' NDCG@5 on ViDoRe v1-v3 while encoding queries up to 26x faster; under identical training OTW matches score distillation with zero page encoding and 12.6x less cached teacher data read.
Related event: ColNanoVDR Distills Multi-Vector Doc Retrieval Without Encoding Documents(2 posts)→
More from Research
- NVIDIA's LSPD brings RL tricks to policy distillation, cutting rollouts by 75% — nvidia · 2026-09-29
- Swapping matmul for associative-algebra layers boosts 110M LM throughput 7.8% — Ilya Koziev · 2026-09-29
- SMAT: merge-aware training lifts merged model scores up to 2.16 with <2% overhead — PolyUHK · 2026-09-29
- KernelZero-7B co-evolution beats Claude 4.5 Sonnet on CUDA kernel generation — Changxin Ke · 2026-09-29
- Frozen-base 34M logit correction module fixes 53.3% of Gemma errors losslessly — eulogik · 2026-09-29
- 6.5M-param NanoForecast beats 200M TimesFM on ETT after pipeline fixes cut MASE 43.8% — eulogik · 2026-09-29