Vector search explained: encode, normalize, compare, rank—and why ANN wins at scale
techNmak · x · 2026-10-02
A clear walkthrough of embedding-based vector search, broken into four operations: encode, normalize, compare, rank.
- Documents pass through an embedding model (all-MiniLM-L6-v2, 384-dim in the example); token reps are mean-pooled excluding padding, L2-normalized, and stored alongside passage IDs.
- Queries go through the same pipeline, so queries and docs share one embedding space.
- With unit-norm vectors, cosine similarity reduces to a dot product: score, sort, return top-k.
- At scale, exact scan cost grows linearly—one query over 1M×384 vectors means 384M coordinate multiplications—so systems use approximate nearest-neighbor indexes (HNSW, IVF) that trade some exactness for speed.
- Key caveat: vector similarity measures embedding-space closeness; it doesn't prove a retrieved passage is correct, complete, or relevant.
More from Research
- OpenAI safety VP Lilian Weng shares her 7-step process for writing research blog posts — SinclairWang1 · 2026-10-02
- Arena.ai Launches HarnessTax: Quantifying How Much the Harness Matters for Coding Agents — solyarisoftware · 2026-10-02
- Berkeley paper: LLMs know your preference changed but still use the old one — rohanpaul_ai · 2026-10-02
- Frozen model, evolving harness: ModularRSI lifts Terminal-Bench 2.0 from 47.57 to 52.43 — jiqizhixin · 2026-10-02
- 30,000 paired QR-code illusions open-sourced with multi-decoder checks and robustness scores — 1roOt · 2026-10-02
- Google's Diffusion Controller: the 90% win rate and gray-box access refer to different setups — Crescitaly · 2026-10-02