Running 35B Model on RTX 3090: Tuning MoE Offload Spikes Processing Speed by 2.36x
Longjumping-Music638 · reddit · 2026-08-06
While running a Qwen3.6-35B-A3B Q6 model on an RTX 3090 (24GB), a developer freed up VRAM by offloading certain MoE expert layers to the CPU. This allowed the batch size to increase from 512 to 1024, boosting prompt processing speed from 564 tok/s to 1330 tok/s—a 2.36x improvement—without sacrificing generation speed.
The author provided full llama-bench reproduction commands and detailed hardware specs. They also revealed using an evolutionary search algorithm (LEVI) to find the optimal configuration automatically. The developer is now asking the community for diverse hardware setups to test the generalizability of this optimization.
More from Infra
- Making the Rust Compiler 3x Faster Could Save Hundreds of Millions Annually — doodlestein · 2026-08-06
- Alphabet Seeks $25B in New Bonds, Raising Over $150B This Year — firstadopter · 2026-08-06
- Jensen Huang Calls Algorithms the New Asset Class; Space Data Center Startup Starcloud Hits $1B Valuation — santoshpanda · 2026-08-06
- Hugging Face Hits Record 4PB Weekly Uploads as AI Treats Humans as Storage — vanstriendaniel · 2026-08-06
- Optimizing DeepSeek on RTX 3090: 128K Context Inference Benchmarks — Ok_Ninja7526 · 2026-08-06
- Run Multiple Models on One GPU: SIE Cuts Self-Hosting Costs 75% — Roger_M_Taylor · 2026-08-06