DDR4 + 7900XTX Runs Qwen3-Next at 45-50 tok/s via Strata, Double llama.cpp Speed

EmPips · reddit · 2026-10-03

A Reddit user reports running Qwen3-Next with IQ3XXS weights (80GB) via Strata on slow DDR4 + a 7900XTX, hitting a stable 45-50 tok/s — roughly double llama.cpp's tuned 22.5 tok/s on the same rig. Similar results reported with 12/16GB cards, faster on DDR5; quality reportedly superior (avoid Q2 weights). If a 27B doesn't fit well, this is a viable alternative — and the author suggests asking an LLM to set it up for your specs.

Original post →

More from Infra

Infra channel →