On-prem LLM stack: the leaks you miss are OCR parsing and hosted eval judges

lucasbennett_1 · reddit · 2026-09-30

Running models locally is the easy part of an on-prem LLM stack, the author argues — the real risk is third-party integrations: one stray HTTP call to a cloud API, hosted eval judge, tracing SaaS or embedding endpoint defeats the whole setup.

Already local: llama.cpp/vLLM/Ollama for runtime, Qwen or Llama models depending on hardware, pgvector or Qdrant for vectors.

Commonly unnoticed leaks:

The author invites others to share fully local end-to-end stacks.

Original post →

More from Infra

Infra channel →