Skip big-provider AI on your phone: llama.cpp + OpenWebUI + Tailscale local setup
GodComplecs · reddit · 2026-10-02
A Redditor shared a simple setup for using local LLMs from your phone instead of big providers:
- Run llama.cpp backend (0.0.0.0:8080) + OpenWebUI (0.0.0.0:8081) on your PC, with local model search enabled
- Use Tailscale to reach your home-hosted models from the phone
- The "secret sauce" for 24GB VRAM: run unsloth's Qwen 3.6 in instruct/non-thinking mode with proper settings for a near-cloud experience
- The author says this fully replaces Google AI Mode etc.; terminal tools like opencode can also route through it (though he hasn't done much agentic stuff)
More from Infra
- US egocentric cleaning data (100s of hrs/week capacity) can't even sell at break-even pricing — paigeinsf · 2026-10-02
- Buyers now fill out export control declarations when purchasing RTX 5090s in stores — blelbach · 2026-10-02
- MachGen pushes MiniMax H3 past its 15s cap with 30-second continuous video — MiniMax_AI · 2026-10-02
- Redditor builds fully local LLM-powered radio site on two DGX Sparks and a 5090 — jwhh91 · 2026-10-02
- VC quip: many neoclouds are closer to 95% than five nines of reliability — saranormous · 2026-10-02
- Report: lenders demand up to 25% collateral from Nvidia as GPU-backed loans wobble — GaryMarcus · 2026-10-02