How to size a local RAG stack for 50 users, from OCR to reranking
InternationalGap3698 · reddit · 2026-08-04
- The poster is building an internal RAG knowledge base for about 50 users and wants to run it locally for privacy and cost reasons.
- They ask for model recommendations across three pieces of the pipeline: OCR for scanned documents, an embedding model for semantic search, and a reranker for better retrieval quality.
- They also want guidance on VM sizing — CPU, RAM, storage, and whether a GPU is needed — plus broader architecture advice for a system where not all users will be active at once.
More from Infra
- Databricks reaches a $4B revenue run-rate and nearly $21.8B raised, says Infra Play — thedealdirector · 2026-08-04
- AI index steepens 5x after late 2024 as compute shifts from pretraining to inference — ProfBuehlerMIT · 2026-08-04
- Local Benchmarking of DeepSeek V4 Flash Quantizations: Q3 vs Q8 — Spicy_mch4ggis · 2026-08-04
- Bittensor’s Root Reborn revives validators and removes automatic sell pressure — const_reborn · 2026-08-04
- ARPL makes llama.cpp adapt to ARM ISA and core topology at runtime — OpeningTough145 · 2026-08-04
- DeepSeek V4 Flash 0731 hits a ctx_other error in llama.cpp speculative decoding — Ambitious_Fold_2874 · 2026-08-04