Should you pay idle costs for local RAG just to keep batch jobs on the serving process?

Cautious_Bit_8521 · reddit · 2026-09-02

The author questions keeping a local RAG process always-on and sized for both retrieval and occasional corpus work. Serving requires predictable latency, while re-embedding, deduplication, and index rebuilding are bursty and may need different CPU, GPU, or memory profiles. The proposed view is to keep vector databases like Milvus on the serving path and attach separate compute only when corpus jobs run, publishing matched snapshots and versions afterwards.

Original post →

More from Infra

Infra channel →