Xiaohongshu’s HELMSMAN moves billion-scale vector search to all-flash servers
小红书技术REDtech · wechat · 2026-07-23
Xiaohongshu’s engine team presented HELMSMAN at OSDI 2026, a full-flash ANNS system designed to cut the cost of large-scale vector search without giving up online SLA. The system targets search, recommendation, advertising, and other high-QPS pipelines where in-DRAM indexes are becoming too expensive.
HELMSMAN combines a clustering-based ANN design, an ANNS-specific storage stack built on SPDK, learned pruning to predict how many clusters to probe, and a GPU/CPU heterogeneous build pipeline. In production, it reportedly replaced a workload that previously needed about 35,000 CPU cores and 350 TB of DRAM with roughly 40 all-flash servers, cutting hardware cost by more than 90% while keeping millisecond latency and strong tail behavior. The paper also reports 2–16× throughput gains over prior DRAM-SSD ANNS systems and much faster rebuild times for billion-scale indexes.
Related event: Xiaohongshu's HELMSMAN: All-Flash Vector Retrieval(2 posts)→
More from Infra
- OpenRouter agents now out-consume humans as AI usage arrives in three waves — AccBalanced · 2026-09-11
- Nvidia Is Now Core to Every Major Robotaxi Stack at Commercial Scale — pdamodaran · 2026-09-11
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- RunningHub open-sources H3Lightning, speeding up MiniMax H3 video generation 12x — 智东西 · 2026-09-11