Is Strix Halo the closest thing to a dream local LLM box? Unified memory vs GPUs for 20B-32B models
Robert__Sinclair · reddit · 2026-10-05
The author surveys options for running 20B-32B models at Q5/Q6: consumer GPUs lack VRAM, NPUs lack memory, FPGAs are bandwidth-bound, while Strix Halo boxes (GMKtec EVO-X2, Framework Desktop) offer huge unified memory at a premium. His dream spec—48GB+ memory, 500+ GB/s, under $1000—doesn't exist yet. He asks owners whether unified memory beats benchmarks in practice and if they regret not buying used 3090/4090 rigs.
More from Infra
- Thoughtworks engineer tests local models for agentic coding on Apple M3 Max and M5 Pro — bibryam · 2026-10-05
- Muse ships reliability fixes after SEVs left cron jobs and scheduled tasks unrecovered — alexandr_wang · 2026-10-05
- AWS Mistakenly Suspends Account, Wabi Down for 3+ Hours With No Recourse — soleio · 2026-10-05
- 539 tok/s DeepSeek on 4x RTX 6000 — and a call-out that community benchmarks inflate 20-30% — HankYeomans · 2026-10-05
- GLM 5.3 flash on dual DGX Sparks gets 50-90% decode boost with new open recipe — swiebertjee · 2026-10-05
- Qualcomm's Snapdragon to power next-gen AI assistants for Meta and OpenAI — ryanshrout · 2026-10-05