Scaling naive retrieval to large collections forces you to reinvent embeddings, researcher explains
antoine_chaffin · x · 2026-09-29
antoinechaffin offers a first-principles thought experiment on why embeddings exist: try scaling naive retrieval to a truly large collection and you'll find you need to pre-filter it, applying full retrieval only to a small subset using something that runs cheaply (a sublinear scan). Do that, and you've just reinvented embeddings from scratch — semantic vector search is the inevitable engineering answer to large-scale retrieval, not magic.
More from Research
- Google Trends' #1 US region for every query is tiny Cheyenne, Wyoming — likely bot traffic — lilyraynyc · 2026-09-29
- K3-Node launches: a Keras 3-native GNN library with 100% PyG API parity across JAX, torch and TF — fchollet · 2026-09-29
- Columbia team uses agents end-to-end for preprint on p53-hormone receptor genomic grammar — anshulkundaje · 2026-09-29
- New preprint finds VLM OCR attention heads that verbalize far more than text, enabling a logit lens for image tokens — gsarti_ · 2026-09-29
- Kipply breaks down transformer inference arithmetic for H200/B200 in new perf engineering repo — ycombinator · 2026-09-29
- Researchers note input length directly affects activation sparsity, call for sequence-compression methods — antoine_chaffin · 2026-09-29