Qwen dual-GPU inference optimization: 10x prefill speed boost achieved

Comrade_Mugabe · reddit · 2026-08-28

Benchmarking Qwen3.8-Flash-Next on dual RTX 3060s, the author found that llama.cpp's default -sm tensor mode causes MoE weights to fall back to CPU, capping prefill at 36 tps. Switching to -sm layer and tuning -ubatch 2048 increased prefill speed to 400 tps. The post compares ikllama.cpp, noting it lacks this trap and matches speed but consumes 75 GB more RAM. Detailed configs and metrics are provided for local deployment.

Original post →

More from Infra

Infra channel →