Alibaba’s H+ Embedding cuts retrieval vectors by 13.7% while matching token-level quality
_reachsumit · x · 2026-08-04
Alibaba’s H+ Embedding introduces context-dependent phrase units between global vectors and token-level retrieval.
The paper argues that terminology-heavy retrieval needs something more granular than a single embedding but cheaper than full token-level late interaction. H+ Embedding predicts variable-length phrase partitions, keeps uncovered tokens as singletons, and applies importance-guided selection with weighted MaxSim. Across 16 scientific, medical, and bilingual tasks, the phrase branch beats the global branch by 6.91 macro nDCG@10, nearly matches token-level retrieval, and uses 13.7% fewer document vectors. The authors position it as a practical middle ground between compression quality and indexing/scoring cost.
More from Infra
- MiniMax H3 Text-to-Video Successfully Runs on a Single DGX Spark — Scobleizer · 2026-08-04
- Spectrum Acceleration for MiniMax H3 in ComfyUI: Up to 34% Lower Inference Time — marres · 2026-08-04
- Counterpoint Research: DRAM and NAND Memory Prices Expected to Peak in 2027 — SumitGup · 2026-08-04
- Atomic-scale logic circuits built from silicon dangling bonds showcased in new paper — teortaxesTex · 2026-08-04
- How a 5% Difference in Cache Hit Rate Spikes Token Cost 3-4x — Xianbao_QIAN · 2026-08-04
- Essential Features for Production LLM AI Gateways — Fun-Beginning5005 · 2026-08-04