Running a 176B MoE on a 16GB RTX 3080 laptop: TensorSharp beats Strata in end-to-end test

fuzhongkai · reddit · 2026-10-04

The author ran Qwen3.8 Flash Next 176B on a RTX 3080 laptop (16GB VRAM) + 32GB RAM + SSD using their open-source inference engine TensorSharp.

The approach treats SSD not as last-resort swap but coordinates memory/storage tiers around MoE execution via quantization + unified scheduling, keeping the right experts in the right tier at the right time.

Benchmark vs Strata:

Takeaway: for huge sparse MoE models the question isn't "does it fit in RAM/VRAM?" but "how efficiently can the runtime coordinate VRAM, RAM, SSD, caching and expert activation?" The author invites comparisons with llama.cpp and Strata on similar hardware.

Original post →

More from Infra

Infra channel →