Qwen dual-GPU inference optimization: 10x prefill speed boost achieved
Comrade_Mugabe · reddit · 2026-08-28
Benchmarking Qwen3.8-Flash-Next on dual RTX 3060s, the author found that llama.cpp's default -sm tensor mode causes MoE weights to fall back to CPU, capping prefill at 36 tps. Switching to -sm layer and tuning -ubatch 2048 increased prefill speed to 400 tps. The post compares ikllama.cpp, noting it lacks this trap and matches speed but consumes 75 GB more RAM. Detailed configs and metrics are provided for local deployment.
More from Infra
- Deepseek Harness: Solving Agent Tool Contention via DB Monitoring — NirantK · 2026-08-28
- CXMT Net Income Hits RMB 77.6B in 1H26, Up 873% YoY — zephyr_z9 · 2026-08-28
- PCIe bottleneck dilemma: Adding RAM vs. GPU for local AI workloads — dsdt · 2026-08-28
- Forked Ninfer for TP2 to achieve 1M context with 50% throughput boost — Littlepharaoh · 2026-08-28
- Local AI is about data ownership, not cost savings — StewartalsopIII · 2026-08-28
- Nvidia Backs $500B Compute Financing Platform, Sparking Subprime Crisis Comparisons — 创业邦 · 2026-08-28