Late-interaction doc embeddings shrink 255x to 6KB per page, keeping 95%+ accuracy

lateinteraction · x · 2026-09-03

For NeoMME-Retriever-260M on ViDoRe v3, high-resolution late-interaction embeddings average 1.5 MB per page. Token pooling plus asymmetric quantization cut this to 6 kB — a 255x reduction — while retaining more than 95% of baseline nDCG@10.

Original post →

More from Research

Research channel →