Framework Desktop with 128GB unified memory: Qwen 27B-35B beats bigger local models
lebbi · reddit · 2026-09-17
The author shares hands-on experience running local LLMs on a Framework Desktop with AMD Strix Halo and 128GB of unified memory:
- Counterintuitive finding: despite 90GB of spare memory, mid-sized models like Qwen 3.6 35B and Qwen 3.8 27B consistently give the best coding results and end up being the daily drivers;
- The bottleneck is CPU: it can't run multiple agents concurrently at decent speed, so much of that memory sits idle;
- The author asks the community what they actually run on Strix Halo machines and whether they're missing something.
Useful reference for anyone considering big-memory machines for local LLMs: more memory doesn't mean bigger models — speed and CPU often limit you first.
More from Infra
- Running a 14B Model on 16GB RAM: 'My PC Is a Toaster Now' — Aggravating_Site381 · 2026-09-17
- fal engineering head: we'll never pre-train, inference compute is the real moat — jfischoff · 2026-09-17
- Meta extends FlashAttention-4 with MXFP8 on Blackwell, hitting 2.85 PFLOP/s forward — PyTorch · 2026-09-17
- AWS shows NVRx fault-tolerant FSDP training on EKS: sync checkpointing ate up to 40% of wall time — AWS ML Blog · 2026-09-17
- Qwen 3.8 27B NVFP4 benchmarked on 2xV100 with ~400k context — jjusko20 · 2026-09-17
- NVIDIA launches CUDA Rust: two paths to write GPU kernels natively in Rust — blelbach · 2026-09-17