Late-interaction doc embeddings shrink 255x to 6KB per page, keeping 95%+ accuracy
lateinteraction · x · 2026-09-03
For NeoMME-Retriever-260M on ViDoRe v3, high-resolution late-interaction embeddings average 1.5 MB per page. Token pooling plus asymmetric quantization cut this to 6 kB — a 255x reduction — while retaining more than 95% of baseline nDCG@10.
More from Research
- Do LLMs Have Research Taste? Lossfunk Paper Tests It via Future Direction Choice — paraschopra · 2026-09-03
- Action Chunking Boosts Contrastive RL Even in Fully Online RL, Study Finds — ben_eysenbach · 2026-09-03
- Coding models are running out of data — PL researchers propose 'intent computing' as the fix — LingmingZhang · 2026-09-03
- davidad Backs Call to Ban Naive RLVR: 'Everything Should Be Model-Graded' — davidad · 2026-09-03
- Computerphile Deep Dive: How Watermarks Track AI-Generated Content — Computerphile · 2026-09-03
- TrafficLab 3D builds digital-twin traffic visualizations from CCTV footage and Google Maps — tom_doerr · 2026-09-03