Dual RTX 5060 Ti Only Gets 10 t/s on Qwen3.8-Flash-Next, Seeking Config Advice

MkGod · reddit · 2026-09-19

A Reddit user is trying to get usable speeds running Qwen3.8-Flash-Next-UD-Q3KXL via Unsloth Studio (llama-server) on dual RTX 5060 Ti 16GB (32GB total VRAM) with 96GB DDR4.

Tried so far:

Decode is only 9.5–10 t/s with 12–38 t/s prefill. Others squeeze 25–40 t/s from dual 3090s using the new expert-cache PRs (#27861/#28223) and pinned host memory, but he's unsure of the best path on 32GB VRAM + DDR4: drop MTP entirely, bypass mmap with --load-mode none, build a custom llama.cpp with LRU expert cache—or is he fundamentally bandwidth-capped?

Original post →

More from Infra

Infra channel →