DDR4 + 7900XTX Runs Qwen3-Next at 45-50 tok/s via Strata, Double llama.cpp Speed
EmPips · reddit · 2026-10-03
A Reddit user reports running Qwen3-Next with IQ3XXS weights (80GB) via Strata on slow DDR4 + a 7900XTX, hitting a stable 45-50 tok/s — roughly double llama.cpp's tuned 22.5 tok/s on the same rig. Similar results reported with 12/16GB cards, faster on DDR5; quality reportedly superior (avoid Q2 weights). If a 27B doesn't fit well, this is a viable alternative — and the author suggests asking an LLM to set it up for your specs.
More from Infra
- Read-only Postgres replicas won't stop LLM agents from bloating your primary DB: 4 failure modes — EmetInteractive · 2026-10-03
- Cathie Wood: 90% of global data center debt financing lands in US at ~8% effective tax — rohanpaul_ai · 2026-10-03
- DeepSeek open-sources DeepGEMM-Ascend: MIT-licensed kernels for Huawei Ascend 950 — lmoroney · 2026-10-03
- Musk says growth will far exceed forecasts as energy remains the bottleneck — RachelVT42 · 2026-10-03
- Running Qwen3 27B agent on 16GB VRAM: MTP draft trades context and quality for speed — Remarkable_Air_8383 · 2026-10-03
- DwarfStar 4 (ds4) lets you run DeepSeek V4.1, Qwen and GLM locally — yogthos · 2026-10-03