A Reddit user designs a 4-layer local AI homelab with vLLM, LiteLLM, TrueNAS and OPNsense

povedaaqui · reddit · 2026-07-21

A Reddit post asks for feedback on a **4-layer local AI homelab architecture** built around strict separation of concerns: - **Inference layer:** vLLM, llama.cpp, Whisper STT, and a VLM for OCR. - **Tools layer:** Hermes agent orchestration, LiteLLM routing, Playwright, SearXNG, Crawl4AI, pandoc, and yt-dlp, each containerized. - **Storage layer:** TrueNAS on bare metal with ZFS mirror and NFS, plus PostgreSQL + pgvector, Git, and SOPS + Age for configs/secrets. - **Network layer:** OPNsense on a dedicated box, Ubiquiti/MikroTik switching, WireGuard, Caddy + Authelia, and AdGuard Home. The routing model is **local first** with cloud overflow via OpenRouter or fal.ai. The author emphasizes zero fixed subscriptions, open source, ARM64 compatibility, and keeping the firewall physically separate. The post asks what others would change before deployment.

Original post →

More from coding & agent

coding & agent channel →