Slashing RAG Latency From 90s to 4s
ezzeddinabdallah · reddit · 2026-07-18
The author shares a RAG pipeline optimization case. Originally, a single query took 90 seconds, causing users to leave before getting an answer.
The root cause wasn't the model, but the retrieval layer itself:
- Overly heavy embeddings
- Lack of caching
- Compounding repeated calls as the document set grew
After refactoring this layer:
- Response time dropped from 90 seconds to about 4 seconds
- Costs decreased by roughly 95%
- Weaviate was used to rebuild retrieval, specifically to fix accuracy issues with "incorrect retrieved content"
The author concludes that in many AI performance issues, the real bottleneck isn't the model, but the overlooked infrastructure layer.
More from Infra
- Spomin: live KV cache compaction squeezes 500k tokens of context into 180k resident — wgaca2 · 2026-09-11
- PiPNN nearest-neighbor search wins three awards, up to 78x faster index building — khademinori · 2026-09-11
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11
- DeepSeek launches V4.1-Flash with 1M-token context and 4x smaller KV-cache — matlabulous · 2026-09-11
- What Can You Still Run on 8GB VRAM? User Asks for Small Models With Tool Use — riceinmybelly · 2026-09-11
- Spain's hourly 80% renewable matching rules clash as France fast-tracks 700MW sites, UK cuts grid queues — eherrerosj · 2026-09-11