Optimized LI index smaller, beats Qwen3-8B by 9 points

tomaarsen · x · 2026-08-26

By pruning to 42% of vectors, the index shrinks to 1.45 GB at 86.4 NDCG@10. In contrast, Qwen3-Embedding-8B's fp16 index for the same corpus is 1.64 GB at 77.5. This proves that a well-configured Late Interaction index can be smaller and significantly outperform dense baselines.

Related event: 1-bit residuals shrink multi-vector indexes 13x(2 posts)→

Original post →

More from Infra

Infra channel →