Optimal parameter settings for local deployment of Qwen3.8 Flash Next shared
Motor_Ad16 · reddit · 2026-08-28
A user shared specific parameter configurations for running Qwen3.8 Flash Next on a dual RTX 3090 and Threadripper Pro 3955 setup. Deployed via Llama CPP in an Ubuntu VM with flash-attn enabled and specific batch-size and ctx-size settings. Benchmarks show a text generation speed (TG) of 15-19 t/s and processing speed (PP) of 550 t/s. The author seeks advice on further optimization.
More from Infra
- Local AI is about data ownership, not cost savings — StewartalsopIII · 2026-08-28
- RTX 3060 12GB: The unsung hero of local AI with 24GB VRAM and 30 t/s — I_Play_Zed · 2026-08-28
- Alibaba Open Sources Qwen3.8-Flash: Undercuts DeepSeek, Runs 1M Context on 4090 — 量子位 · 2026-08-28
- Using Langfuse traces to autonomously analyze and improve agent workflows — NielsRogge · 2026-08-28
- Nvidia arranged $500B in AI infra financing, guaranteeing $105B for OpenAI — VraserX · 2026-08-28
- Micron: HBM Requires Three Times More Wafer Area Than DDR5 — FullstackSensei · 2026-08-28