On-prem LLM stack: the leaks you miss are OCR parsing and hosted eval judges
lucasbennett_1 · reddit · 2026-09-30
Running models locally is the easy part of an on-prem LLM stack, the author argues — the real risk is third-party integrations: one stray HTTP call to a cloud API, hosted eval judge, tracing SaaS or embedding endpoint defeats the whole setup.
Already local: llama.cpp/vLLM/Ollama for runtime, Qwen or Llama models depending on hardware, pgvector or Qdrant for vectors.
Commonly unnoticed leaks:
- Ingestion: PDF/scan parsing before chunking often reaches for a cloud parser — LiteParse-style open-source parsers or HF alternatives keep it local
- Eval: hosted judges and tracing dashboards receive prompts and outputs — swap in a local score set, a local judge model, and self-hosted tracing
The author invites others to share fully local end-to-end stacks.
More from Infra
- 5 ways to cut LLM costs without changing models: optimize tokens, caching and calls — goyalshaliniuk · 2026-09-30
- DeepSeek Now Training on Huawei Ascend 950 Chips — WebAssemblyMan · 2026-09-30
- FreeToken: open-source engine runs 290B+ MoE models locally on consumer hardware — tom_doerr · 2026-09-30
- EnerTune at SOSP'26 cuts LLM serving energy 1.4-2.3x vs SOTA systems — tianyin_xu · 2026-09-30
- SOSP'26 paper proposes energy-conscious GPU sharing for inference serving, beyond utilization — tianyin_xu · 2026-09-30
- SOSP'26 papers: AgileLog for isolated AI-agent data streams, CXL-LSM for CXL shared memory — tianyin_xu · 2026-09-30