Dual RTX 5060 Ti Only Gets 10 t/s on Qwen3.8-Flash-Next, Seeking Config Advice
MkGod · reddit · 2026-09-19
A Reddit user is trying to get usable speeds running Qwen3.8-Flash-Next-UD-Q3KXL via Unsloth Studio (llama-server) on dual RTX 5060 Ti 16GB (32GB total VRAM) with 96GB DDR4.
Tried so far:
- CPU-offloaded experts via --cpu-moe -ngl 99 -b 1024 -ub 512 -fa on
- 65k context, q40 KV cache
- MTP speculative decoding (87% draft acceptance)
Decode is only 9.5–10 t/s with 12–38 t/s prefill. Others squeeze 25–40 t/s from dual 3090s using the new expert-cache PRs (#27861/#28223) and pinned host memory, but he's unsure of the best path on 32GB VRAM + DDR4: drop MTP entirely, bypass mmap with --load-mode none, build a custom llama.cpp with LRU expert cache—or is he fundamentally bandwidth-capped?
More from Infra
- GPU host warns: renter exploited his rig for attacks, Clore.AI blocked him for reporting it — anomaly256 · 2026-09-19
- 3M paid $10.3B to quit PFAS — AI data centers just made it a growth market again — aakashgupta · 2026-09-19
- $250 of modded mining cards, 30GB VRAM: old i7 PC runs Qwen at 30 tok/s with patched drivers — HFq_Dev · 2026-09-19
- Reading a Pretraining Run: A P0/P1/P2 Metric System for Monitoring LLM Pretraining — SonglinYang4 · 2026-09-19
- Hacking open-source model behavior with sglang's scoring endpoint, no fine-tuning needed — BLUECOW009 · 2026-09-19
- Distilling DeepSeek V4 Flash to a 4B model on DGX Spark: 26 hours, 22ms per judgment — Dan_Jeffries1 · 2026-09-19