Running Qwen3.8 Flash Next on dual RTX 3090: full llama.cpp config shared for tuning

ChopSticksPlease · reddit · 2026-09-12

A user shares their llama.cpp setup for Qwen3.8 Flash Next on dual RTX 3090s (48GB VRAM) + 128GB DDR4 + 40-core Xeon under Proxmox: 131072 ctx, q80 KV cache, flash attention, -ts 26,10, -ncmoe 26, and per-layer tensor overrides. Current results: PP 130-200 tps, TG 15 tps average; asking for tuning advice.

Original post →

More from Infra

Infra channel →