Struggling to match Strata speeds running Qwen3.8-Flash-Next in vanilla llama.cpp on 2x RTX 3060
Dreeew84 · reddit · 2026-10-12
A Redditor gets only 15 tok/s running Qwen3.8-Flash-Next IQ3XXS in vanilla llama.cpp on 2x RTX 3060 vs Strata's 35-50 tok/s. Issues: can't find the 800MB mtp q20 drafter Strata uses (only a 3.5x larger Q4KM that slows decode), and tensor split fails with cudaMalloc errors as llama treats the cards as one 24GB unit. Asking for working recipes.
More from Infra
- $2 ESP32 board runs Pi-hole-style DNS blocker with 140K domains in 0.7MB — LinusEkenstam · 2026-10-12
- Engineer joins NVIDIA's Groq LPU compilers team, working on multi-chip partitioning — blelbach · 2026-10-12
- Linus Ekenstam wants nothing less than a 100B-param model running on your phone — LinusEkenstam · 2026-10-12
- AI buildout to cost $10.3 trillion to finance through 2032, topping all prior US investment booms — KyeGomezB · 2026-10-12
- Fireworks: open models plus fine-tuning match closed ones — Cursor gets 13x faster inference — AI Engineer · 2026-10-12
- Running a 456GB model on 192GB VRAM: offloaded inference hits 60-125 tok/s with 1M context — HankYeomans · 2026-10-12