Struggling to match Strata speeds running Qwen3.8-Flash-Next in vanilla llama.cpp on 2x RTX 3060

Dreeew84 · reddit · 2026-10-12

A Redditor gets only 15 tok/s running Qwen3.8-Flash-Next IQ3XXS in vanilla llama.cpp on 2x RTX 3060 vs Strata's 35-50 tok/s. Issues: can't find the 800MB mtp q20 drafter Strata uses (only a 3.5x larger Q4KM that slows decode), and tensor split fails with cudaMalloc errors as llama treats the cards as one 24GB unit. Asking for working recipes.

Original post →

More from Infra

Infra channel →