Should you pay idle costs for local RAG just to keep batch jobs on the serving process?
Cautious_Bit_8521 · reddit · 2026-09-02
The author questions keeping a local RAG process always-on and sized for both retrieval and occasional corpus work. Serving requires predictable latency, while re-embedding, deduplication, and index rebuilding are bursty and may need different CPU, GPU, or memory profiles. The proposed view is to keep vector databases like Milvus on the serving path and attach separate compute only when corpus jobs run, publishing matched snapshots and versions afterwards.
More from Infra
- Cloudflare Agents emit OpenTelemetry traces, route directly to Braintrust for evals — ritakozlov · 2026-09-02
- Rabbi's take on DC moratorium: Using bans as leverage for environmental and labor concessions — joshua_saxe · 2026-09-02
- Intel exec: AI era security requires silicon-level design, not afterthoughts — BenBajarin · 2026-09-02
- PyTorch 2.14 released with 2,995 commits from 487 contributors — PyTorch · 2026-09-02
- Exllamav3 benchmarks: 700tk/s on 8x3090 setup — Leflakk · 2026-09-02
- GLM-5.3 Model Gets GGUF Quantization Release for Edge Deployment — unsloth · 2026-09-02