Dev turns a Pixel 10 Pro XL into an offline OpenAI-compatible server running Gemma 4 E4B at ~11 tok/s
VerityAISolutions · reddit · 2026-10-03
A developer open-sourced PixelUnlockGPU, an Android app that turns a Pixel 10 Pro XL (Tensor G5, 16GB) into an OpenAI-compatible local HTTP server running Gemma 4 E4B instruct (Apache 2.0, 2.97GB) fully offline via LiteRT-LM's GPU backend.
- Measured 11 tok/s steady-state decode; KV-prefix reuse across turns gives 1.7s follow-ups, so stateless clients like TypingMind never re-prefill history
- Three access modes: loopback-only, raw LAN, or Tailscale with WireGuard end-to-end encryption
- Honest limitations: NPU path aborts on stock Tensor G5 firmware (GPU is the shipping backend), no GPU temp sensor access, no per-request stop API, token usage only estimated (4 chars/token)
- Chose the official LiteRT-LM GPU path, an always-on Ktor/Netty service, and the OpenAI wire over llama.cpp/ollama
Repo and v0.1.0/v0.1.1 APKs are on GitHub; contributions welcome.
More from coding & agent
- "I Gave My Agent /tmp and It Moved In": Kevin Kern Shares Agent Humor — kevinkern · 2026-10-03
- Running Grok Bot on multiple Macs at once to scale your AI bot army — Baconbrix · 2026-10-03
- Real-world agent test: one nails church revenue analysis one-shot — MikkoH · 2026-10-03
- Agent workflow: triage Slack bugs, auto-open draft PRs when info suffices — gabrielchua · 2026-10-03
- Meta Muse hits 5M downloads as AI feedback loops could build a fast moat — RihardJarc · 2026-10-03
- Context.dev launches Highlights web-scraping API built for AI agents and LLMs — DanielLockyer · 2026-10-03