Local LLM setup: Strata lets a 7900XTX + 64GB RAM run Qwen Flash at 60 tok/s at 250K context
soyalemujica · reddit · 2026-10-03
A Reddit user reports running Qwen Flash via Strata on a 7900XTX (24GB VRAM) + 64GB DDR5 setup, hitting 60 tokens per second even at 250K context and q8 precision.
They claim it outperforms the dense 27B model they dropped — better instruction following, more reliable planning, and capable of running frontend/backend test loops — while eliminating OOM errors and freeing the PC for gaming.
Notable details: a 6GB VRAM reserve keeps Windows 11 usable, and Windows performance now matches Linux. A practical data point for local MoE deployments.
More from Infra
- Dev platform cuts prices in half, making GitHub Action runners 12x cheaper — aniketmaurya · 2026-10-03
- Free 5% speedup: enable CustomAllReduce on SM120 to push tensor parallelism from TP=2 to TP=4 — TheZachMueller · 2026-10-03
- Hand-written Blackwell GEMM in CuTe DSL hits 1401 TFLOP/s, about 97% of cuBLAS on B200 — retr0jirachi · 2026-10-03
- David Patterson: solar's near-vertical cost curve will power the singularity — davidpattersonx · 2026-10-03
- DGX Spark shortage derails $8,800 donation plan as buyer can't find stock — cyrus_zei · 2026-10-03
- Perplexity to vertically integrate agentic infra on NVIDIA's Vera CPU, ditches x86 — AravSrinivas · 2026-10-03