Reddit debates the cheapest practical way to run K3 locally
ylchao · reddit · 2026-07-28
A Reddit thread asks what the cheapest viable way would be to run K3 locally, and proposes a wide range of hardware setups and hacks.
Ideas mentioned include DGX Spark or Strix Halo clusters, Optane persistent memory with GPUs, Mac Studio clusters, Orange Pi 6 clusters, SSD streaming plus GPUs, DDR3 + ConnectX-5 RDMA clients, two DGX Stations, and even Power10 systems.
The thread is essentially about the infrastructure economics of hosting a large model locally.
More from Infra
- MLA-based KV cache costs 12 GB per million tokens, with KDA state at 230 MB BF16 — zephyr_z9 · 2026-07-28
- A 70B GGUF model stalls on an AMD R9700 as VRAM hits 30 GB but RAM stays flat — Developer-Y · 2026-07-28
- K3 hit a usage-limit exploit on Higgsfield within a day of launch — Mediocre-Witness-778 · 2026-07-28
- TRELLIS.2 INT8 ConvRot now runs natively in ComfyUI on AMD ROCm — DrBearJ3w · 2026-07-28
- AI data-center buildout is reshaping grid incentives and reserve power — kleffew94 · 2026-07-28
- The Verge says Moonshot’s open Kimi K3 could undercut closed U.S. AI models — The Verge AI · 2026-07-28