The curse of 64GB RAM: Strata pushes local Qwen3.8-Flash-Next to 60 t/s but hogs system memory
Cautious_Chicken_604 · reddit · 2026-10-04
A local LLM enthusiast with an R9700 + RTX 5060 Ti + 64GB DDR5 shares a tradeoff story: the R9700 runs Qwen3.8-27B Q6 as a daily driver (35 t/s) while the 5060 Ti handles ComfyUI image/video inference.
- Running Qwen3.8-Flash-Next's 93GB IQ4XS quant on both cards via Vulkan only hit 15 t/s — unusable.
- With Strata running the 70GB IQ3XXS quant, llama.cpp on system RAM + R9700 hit 21 t/s; Strata reached 60 t/s, finally daily-drivable.
- The catch: QFN loaded in Strata pushes system RAM to 96%, making memory-hungry ComfyUI workloads like Minimax H3 impossible to run simultaneously.
The author is torn between finally running the big model and losing his media-generation workflow, with a 128GB RAM upgrade out of budget for now.
More from Infra
- BF16 rounding breaks a conservation law, blowing up FlashAttention gradients late in training — HongyiWang10 · 2026-10-04
- Getting PyTorch CUDA training running on BC-250 boards, captured as an image — redfoxkiller · 2026-10-04
- Huawei 950 super-node claims seamless scaling from 550B/1.6T up to 10T-class models — teortaxesTex · 2026-10-04
- Huawei: 1000+ Ascend 910C supernodes deployed, 40+ LLMs natively pretrained — teortaxesTex · 2026-10-04
- Huawei's Ascend 950DT TDP hits 950W per card; B200 ~4.4x denser in FP8/W — teortaxesTex · 2026-10-04
- China Unicom's Heterogeneous Grouped MoE Cuts Parameters 20%, Wins ACL 2026 Spot — 量子位 · 2026-10-04