awesome-multivector-retrieval now tracks 50+ datasets across 17 corpora and 8 encoders
CShorten30 · x · 2026-09-29
The community-maintained awesome-multivector-retrieval repo keeps growing: after Silvio Martinico published 9 multivector datasets last week, Robro612 added 40+ more on Hugging Face. The repo now lists 50+ ready-to-use datasets spanning 17 corpora and 8 encoders.
The repo is an annotated resource list for late-interaction multivector retrieval, covering foundational models (ColBERT, COIL), compression and token pruning, multimodal/vision models, indexing and search algorithms, scoring kernels, software libraries and training/inference frameworks, model checkpoints, and datasets/encodings such as NFCorpus, SciFact, MS MARCO and ViDoRe.
More from Research
- Statistician mocks frequentists for patching one flaw and ignoring the rest — RexDouglass · 2026-09-29
- Pre-registration is just year-old priors, statistician argues in Bayesian-frequentist spat — RexDouglass · 2026-09-29
- 300M-param image model trained for €5000 claims SD 1.5-level performance — incorporo · 2026-09-29
- Cheap verifiers match costly ones in LLM post-training, saving up to 99.7% of grading cost — iScienceLuvr · 2026-09-29
- Telescopic LM trains one model valid at every depth, cutting quality-budget area 43% — iScienceLuvr · 2026-09-29
- ROFT: fine-tuning only on self-explanations matches GRPO on SWE-bench without RL — iScienceLuvr · 2026-09-29