I stopped renting GPUs: two boxes and an iPhone app run my whole AI stack
EAccelerate_42 · x · 2026-09-05
The author ditched rented GPUs for two small boxes in a closet (121 GiB RAM + 48 GB VRAM) controlled via an iPhone app. Local stack includes DeepSeek V4 Flash (384K context), Qwen3.8 27B, GLM-5.3 Flash, Qwen-Image with editing, and MiniMax H3 video. Access is Tailscale-only, with zero tokens billed.
More from Infra
- ESP32 voice assistant: AIMET AdaRound cuts wake-word model 8x without the 3.7% accuracy hit of naive 4-bit — carrycooldude · 2026-09-05
- Dev boosts GLM 5.2 TPS on a B300 and swaps it into Claude Code in place of Anthropic models — abhijithneil · 2026-09-05
- What forces LLM teams to optimize inference when going from MVP to production? — Ok_Philosophy_4031 · 2026-09-05
- Leaker claims GPT-6 Astra beats Claude Fable 5.1 at metaprompting — whoiskatrin · 2026-09-05
- Local Qwen 27B vibecodes a playable Godot dungeon game in just 4 prompts — jacek2023 · 2026-09-05
- Open-sourcing nopasswd-sudo: time-boxed passwordless sudo for agents, built by a local 124B model — max_paperclips · 2026-09-05