A Reddit user designs a 4-layer local AI homelab with vLLM, LiteLLM, TrueNAS and OPNsense
povedaaqui · reddit · 2026-07-21
A Reddit post asks for feedback on a **4-layer local AI homelab architecture** built around strict separation of concerns: - **Inference layer:** vLLM, llama.cpp, Whisper STT, and a VLM for OCR. - **Tools layer:** Hermes agent orchestration, LiteLLM routing, Playwright, SearXNG, Crawl4AI, pandoc, and yt-dlp, each containerized. - **Storage layer:** TrueNAS on bare metal with ZFS mirror and NFS, plus PostgreSQL + pgvector, Git, and SOPS + Age for configs/secrets. - **Network layer:** OPNsense on a dedicated box, Ubiquiti/MikroTik switching, WireGuard, Caddy + Authelia, and AdGuard Home. The routing model is **local first** with cloud overflow via OpenRouter or fal.ai. The author emphasizes zero fixed subscriptions, open source, ARM64 compatibility, and keeping the firewall physically separate. The post asks what others would change before deployment.
More from coding & agent
- X post asks whether Cursor Composer, built on Kimi models, would also be banned — max_paperclips · 2026-07-21
- A developer’s Codex usage is draining pooled enterprise credits at a small company — Distinct_Relation_62 · 2026-07-21
- Qwen Code ships cua-driver-rs 0.7.3 with relative coordinates and MCP filtering — github-actions[bot] · 2026-07-21
- Matt Pocock says every new codebase turns legacy within days — mattpocockuk · 2026-07-21
- Meta and Unity link AI workflows to Quest development across setup, input and validation — Vjeux · 2026-07-21
- AI Engineer World’s Fair spotlights Kids Day with 87 children learning to code — steveonjava · 2026-07-21