Benchmark: Qwen 3.8 hits 45 tps with 1M context on 3080Ti + Strix Halo

TrifleHopeful5418 · reddit · 2026-08-18

The author shared performance data for running Qwen 3.8 on an AMD Strix Halo (128GB) + eGPU 3080Ti (12GB) setup:

While slower than a 5090 setup, there is ample memory remaining to load Embedding and Reranker models simultaneously. The author plans to compare this against Qwen 3.6-35B running on 4x3090 (450 tps).

Original post →

More from Infra

Infra channel →