Running a 176B Qwen MoE on a 16GB RTX 3080 laptop — 4x faster than Strata end-to-end

fuzhongkai · reddit · 2026-10-04

The author ran Qwen3.8 Flash Next 176B on an RTX 3080 Laptop (16GB VRAM) + 32GB RAM + SSD using TensorSharp, his open-source local inference engine — no 128GB RAM workstation or multi-GPU setup needed.

The approach: quantization + MoE-aware unified scheduling across cache, VRAM, system RAM, and SSD. Instead of treating the SSD as last-resort swap, TensorSharp coordinates memory tiers around MoE execution to keep the right experts/data in the right tier at the right time.

Benchmark vs Strata on the same hardware:

Takeaway: for huge sparse MoE models the question isn't "do I have enough RAM/VRAM to fit the model?" but "how efficiently can the runtime coordinate VRAM, RAM, SSD, caching, and expert activation?"

Related event: TensorSharp runs 176B MoE model on a 16GB RTX 3080 laptop(2 posts)→

Original post →

More from Infra

Infra channel →