Perplexity explains its embedding search: batch vs online workloads in one vector space

perplexity_ai · x · 2026-09-05

Perplexity explains that queries and documents are embedded into one vector space and searched by nearest vectors, creating two distinct workloads: throughput-focused bulk batch embedding for indexing and scoring, and latency-focused per-query online embedding for live search—each requiring different infrastructure optimizations.

Related event: Perplexity Reveals Its In-House Embedding Inference Stack(8 posts)→

Original post →

More from Infra

Infra channel →