Strix Halo + 3090 Ti Pushes Qwen Flash-Next to 84 tok/s via Deep Optimization
TrifleHopeful5418 · reddit · 2026-09-02
A user achieved 84 tok/s aggregate decode on Qwen3.8-Flash-Next (104GB) using a hybrid setup: AMD Strix Halo (395, 128GB) + RTX 3090 Ti eGPU. Through seven specific optimizations—including fixing DeltaNet snapshot rollback over PCIe, optimizing iGPU buffer reads, and disabling multi-stream speculation—throughput jumped from 22.2 tok/s. In HumanEval+ testing, this local setup had a median task time of 22.5s (0.4x of a remote dual-3090 vLLM cluster) and passed 155/164 problems, missing just one more than the remote baseline.
More from Infra
- Zyphra open-sources PUFFER, a CPU-based deduplication system 35x faster — bclavie · 2026-09-02
- Dell reports record $47B quarterly revenue, raises outlook on AI infrastructure demand — Polymarket · 2026-09-02
- Nvidia cuts Rubin Ultra memory to 192GB as HBM costs soar to 40% of TCO — rwang07 · 2026-09-02
- SpaceX data center team shakeup: Musk replaces leaders with rocket, satellite internet execs — kyliebytes · 2026-09-02
- Seeking Datacenter-Grade OCS — jwt0625 · 2026-09-02
- GB10 price hike sparks debate: Is Mac Studio the best value for compute? — geekender · 2026-09-02