DIY $3k Strix Halo + unlocked CMP 170HX rig hits 2000 tok/s aggregate, replacing cloud LLMs
Edenar · reddit · 2026-10-10
A detailed writeup of a $3k local LLM rig: a 128GB Framework Strix Halo desktop plus a $500 CMP 170HX unlocked to 64GB via the open-source cmpunlocker tool, connected over a USB4 eGPU dock (pcie2x4).
- Setup: Strix Halo runs qwen 3.8 flash next as the main agent (1600-1800 tok/s prefill, 60 tok/s decode); the CMP 170HX concurrently serves two qwen 3.8 27b subagents with 200k bf16 context each, for 2000 tok/s pp and 150-300 tok/s tg aggregate at low context. A real 4-hour session (7.2M in / 1.1M out tokens) averaged 1747 tok/s pp, 100 tok/s tg.
- Workflow: Pi agent architecture where the main model reviews subagent output, with websearch/python tools for DevOps work (CI/CD, containers). The author rates it above Sonnet 4.7, below Opus; 1/100 tool calls fail and need retries, and a day's work now takes 20 minutes.
- Costs: fully loaded draw 400W, cloud usage dropped entirely; replicating today would cost $5-6k. Mixing AMD gfx1151 and CUDA SM80 for larger models proved impractical even in llama.cpp. Streaming ngram tables from an Optane P5800X showed no gain over a regular PCIe 4 NVMe.
More from coding & agent
- Testing proactive AI agents on my email and calendar: dot caught a date mixup, Muse is noisy — hazelcough · 2026-10-10
- Personal AI agents are stickier than you think: one now learns work habits and builds its own CRM — thisiskp_ · 2026-10-10
- Creator finds hand-tweaking generative models faster than prompts, sees room beyond text UIs — keenanisalive · 2026-10-10
- "Do better!" prompting stalls fast; even top VLMs understand images unevenly — keenanisalive · 2026-10-10
- Telling an LLM to "believe in yourself" helps it write 3D SDF models, but not enough — keenanisalive · 2026-10-10
- This 3D dragon is 27KB of LLM-generated GLSL, not a mesh or NeRF — keenanisalive · 2026-10-10