Microsoft's EigenLI Compresses ColBERT Representations, Beats MUVERA Baseline
_reachsumit · x · 2026-09-09
Microsoft researchers present EigenLI, a training-free spectral method that compresses ColBERT-style late-interaction representations. Key insight: document token embeddings concentrate in a low-dimensional subspace that preserves most retrieval signal, so each document's dominant eigendirections can build a reduced interaction representation. k-EigenLI with k≤32 outperforms k-means and Ward clustering pooling on ColBERTv2 and AnswerAI-ColBERT-small, though GTE-ModernColBERT favors clustering at k=32. The same construction also yields EigenLI-SV, an ANN-compatible single-vector representation that consistently beats MUVERA-like surrogates across datasets and models.
More from Research
- Timothy Duff's ECCV 2026 SfM-DL workshop slides on algebraic optimality for minimal solvers — ducha_aiki · 2026-09-09
- Drop a fixed batch proportion instead of per-sample tokens: capi author shares training trick — giffmana · 2026-09-09
- Adding Greek to a Cosmos3 VLA policy: bilingual training helps but lags far behind English — KIEFERSA · 2026-09-09
- Transformers encode a partner's expertise early but only act on it in later layers — Mika Okamoto · 2026-09-09
- Cadence uses a time-series foundation model for error-bounded lossy compression of demand data — Roberto Tacconelli · 2026-09-09
- Fourth Perception Test Challenge at ECCV 2026 pushes multimodal models on city-scale spatial intelligence — AjdDavison · 2026-09-09