Is one fine-tuned small model per tenant a viable alternative to RAG?
alexrada · reddit · 2026-09-29
A developer is weighing a multi-tenant alternative to RAG: fine-tune a small model (SLM) per client so the model itself holds the client's data. He knows it's technically possible but wants real-world input on efficiency, when it's worth doing, and costs. Early-stage research question with no conclusions or numbers yet.
More from Infra
- BioNeMo team boosts Mixtral-8x7B training throughput 2.21x vs HF BF16 baseline — AllThingsApx · 2026-09-29
- DeepSeek's elastic compute team is hiring heavily, shares sandbox infra for large-scale agent training — teortaxesTex · 2026-09-29
- Developer slams third-party inference providers: Gemini up 10x, Luna 15s latency — julianharris · 2026-09-29
- Nereus: adaptive parallelism boosts 8B PPO throughput up to 7.27x over OpenRLHF — Songlin Jiang · 2026-09-29
- Data center water use isn't about total volume, it's who runs out of local freshwater first — AryHHAry · 2026-09-29
- cache-proxy Ships Exact and Semantic LLM Caching with x402 USDC Payments on Base — modelcontextprotocol · 2026-09-29