Running a 176B MoE on a 16GB RTX 3080 laptop: TensorSharp beats Strata in end-to-end test
fuzhongkai · reddit · 2026-10-04
The author ran Qwen3.8 Flash Next 176B on a RTX 3080 laptop (16GB VRAM) + 32GB RAM + SSD using their open-source inference engine TensorSharp.
The approach treats SSD not as last-resort swap but coordinates memory/storage tiers around MoE execution via quantization + unified scheduling, keeping the right experts in the right tier at the right time.
Benchmark vs Strata:
- Decode: TensorSharp 11.09 tok/s vs 10.24 tok/s
- Whole-process time: 16.54s vs 62.15s
- GPU peak: 14,832 MiB vs 15,729 MiB
- OS working set: 19.74 GiB vs 18.51 GiB
Takeaway: for huge sparse MoE models the question isn't "does it fit in RAM/VRAM?" but "how efficiently can the runtime coordinate VRAM, RAM, SSD, caching and expert activation?" The author invites comparisons with llama.cpp and Strata on similar hardware.
More from Infra
- Samsung: HBM to consume 30% of DRAM wafer capacity by 2027, up from ~20% — Beth_Kindig · 2026-10-04
- Nebius up 160% vs CoreWeave's 10%: the deciding factor isn't revenue or backlog — Beth_Kindig · 2026-10-04
- US Data Center Construction Spending Surged 73% YoY to a Record $85B Annualized — FlorianGallwitz · 2026-10-04
- Developer's talk on running on-device AI praised for great insights — carrycooldude · 2026-10-04
- Cloudflare: pending I/O now keeps Durable Objects alive for agents without a connected client — irvinebroque · 2026-10-04
- Dev ships C-based local inference engine targeting tool calling on low-VRAM hardware — ZenZombie117 · 2026-10-04