Benchmark: Qwen 3.8 hits 45 tps with 1M context on 3080Ti + Strix Halo
TrifleHopeful5418 · reddit · 2026-08-18
The author shared performance data for running Qwen 3.8 on an AMD Strix Halo (128GB) + eGPU 3080Ti (12GB) setup:
- 262K Context (Q6): iGPU only 10 tps; with MTP (n=4) 24 tps; with FastMTP offloaded to 3080Ti 28 tps.
- 262K Context (Q4km): Layer splitting (weights + FastMTP on 3080Ti) 53 tps.
- 1M Context (Q4km): Layer splitting 45 tps.
While slower than a 5090 setup, there is ample memory remaining to load Embedding and Reranker models simultaneously. The author plans to compare this against Qwen 3.6-35B running on 4x3090 (450 tps).
More from Infra
- Cloudflare Durable Objects gets a deployments tab showing real traffic per version — ritakozlov · 2026-08-18
- Magnitude Catalog Profiles Hardware and Recommends Models — Dan_Jeffries1 · 2026-08-18
- Local vs. Sovereign AI: Where Is the Industry Drawing the Line? — rio_ARC · 2026-08-18
- City2Graph library turns geospatial data into spatial graphs — tom_doerr · 2026-08-18
- Interactive Guide: How to Parallelize a Transformer for Training — ezyang · 2026-08-18
- Qwen3.8-27B benchmarks and SGLang high-throughput serving guide — Sam Witteveen · 2026-08-18