Strix Halo + 3090 Ti Pushes Qwen Flash-Next to 84 tok/s via Deep Optimization

TrifleHopeful5418 · reddit · 2026-09-02

A user achieved 84 tok/s aggregate decode on Qwen3.8-Flash-Next (104GB) using a hybrid setup: AMD Strix Halo (395, 128GB) + RTX 3090 Ti eGPU. Through seven specific optimizations—including fixing DeltaNet snapshot rollback over PCIe, optimizing iGPU buffer reads, and disabling multi-stream speculation—throughput jumped from 22.2 tok/s. In HumanEval+ testing, this local setup had a median task time of 22.5s (0.4x of a remote dual-3090 vLLM cluster) and passed 155/164 problems, missing just one more than the remote baseline.

Original post →

More from Infra

Infra channel →