Running a 176B Qwen MoE on a 16GB RTX 3080 laptop — 4x faster than Strata end-to-end
fuzhongkai · reddit · 2026-10-04
The author ran Qwen3.8 Flash Next 176B on an RTX 3080 Laptop (16GB VRAM) + 32GB RAM + SSD using TensorSharp, his open-source local inference engine — no 128GB RAM workstation or multi-GPU setup needed.
The approach: quantization + MoE-aware unified scheduling across cache, VRAM, system RAM, and SSD. Instead of treating the SSD as last-resort swap, TensorSharp coordinates memory tiers around MoE execution to keep the right experts/data in the right tier at the right time.
Benchmark vs Strata on the same hardware:
- Decode throughput: 11.09 vs 10.24 tok/s (close)
- End-to-end time: 16.54s vs 62.15s (4x faster)
- Peak VRAM: 14,832 MiB vs 15,729 MiB
Takeaway: for huge sparse MoE models the question isn't "do I have enough RAM/VRAM to fit the model?" but "how efficiently can the runtime coordinate VRAM, RAM, SSD, caching, and expert activation?"
Related event: TensorSharp runs 176B MoE model on a 16GB RTX 3080 laptop(2 posts)→
More from Infra
- Huawei: 1000+ Ascend 910C supernodes deployed, 40+ LLMs natively pretrained — teortaxesTex · 2026-10-04
- Huawei's Ascend 950DT TDP hits 950W per card; B200 ~4.4x denser in FP8/W — teortaxesTex · 2026-10-04
- The curse of 64GB RAM: Strata pushes local Qwen3.8-Flash-Next to 60 t/s but hogs system memory — Cautious_Chicken_604 · 2026-10-04
- A 5KB pure x86-64 assembly engine runs Gemma-2B at 4.6 tok/s on CPU — tom_tsai28 · 2026-10-04
- Bab, a BLAKE3-Inspired Hash Function Family With Streaming Verification, Goes Open Source — carsonfarmer · 2026-10-04
- Compute Is the New Currency: OpenAI Gave YC Startups $2M in Credits for Equity — AccBalanced · 2026-10-04