Running 35B Model on RTX 3090: Tuning MoE Offload Spikes Processing Speed by 2.36x

Longjumping-Music638 · reddit · 2026-08-06

While running a Qwen3.6-35B-A3B Q6 model on an RTX 3090 (24GB), a developer freed up VRAM by offloading certain MoE expert layers to the CPU. This allowed the batch size to increase from 512 to 1024, boosting prompt processing speed from 564 tok/s to 1330 tok/s—a 2.36x improvement—without sacrificing generation speed.

The author provided full llama-bench reproduction commands and detailed hardware specs. They also revealed using an evolutionary search algorithm (LEVI) to find the optimal configuration automatically. The developer is now asking the community for diverse hardware setups to test the generalizability of this optimization.

Original post →

More from Infra

Infra channel →