Local LLM setup: Strata lets a 7900XTX + 64GB RAM run Qwen Flash at 60 tok/s at 250K context

soyalemujica · reddit · 2026-10-03

A Reddit user reports running Qwen Flash via Strata on a 7900XTX (24GB VRAM) + 64GB DDR5 setup, hitting 60 tokens per second even at 250K context and q8 precision.

They claim it outperforms the dense 27B model they dropped — better instruction following, more reliable planning, and capable of running frontend/backend test loops — while eliminating OOM errors and freeing the PC for gaming.

Notable details: a 6GB VRAM reserve keeps Windows 11 usable, and Windows performance now matches Linux. A practical data point for local MoE deployments.

Original post →

More from Infra

Infra channel →