Optimal parameter settings for local deployment of Qwen3.8 Flash Next shared

Motor_Ad16 · reddit · 2026-08-28

A user shared specific parameter configurations for running Qwen3.8 Flash Next on a dual RTX 3090 and Threadripper Pro 3955 setup. Deployed via Llama CPP in an Ubuntu VM with flash-attn enabled and specific batch-size and ctx-size settings. Benchmarks show a text generation speed (TG) of 15-19 t/s and processing speed (PP) of 550 t/s. The author seeks advice on further optimization.

Original post →

More from Infra

Infra channel →