Running Qwen3.8 Flash Next on dual RTX 3090: full llama.cpp config shared for tuning
ChopSticksPlease · reddit · 2026-09-12
A user shares their llama.cpp setup for Qwen3.8 Flash Next on dual RTX 3090s (48GB VRAM) + 128GB DDR4 + 40-core Xeon under Proxmox: 131072 ctx, q80 KV cache, flash attention, -ts 26,10, -ncmoe 26, and per-layer tensor overrides. Current results: PP 130-200 tps, TG 15 tps average; asking for tuning advice.
More from Infra
- TensorSharp hits 41 tok/s decoding DeepSeek V4.1 Flash on 8× A40 — fuzhongkai · 2026-09-12
- What actually runs AI models at the edge in 2026: Mac mini, DGX Spark, iPhone 17 Pro — MaziyarPanahi · 2026-09-12
- UAE redesigns 5GW AI campus with bunkers and air defenses after Iranian strikes on Gulf cloud facilities — mark_k · 2026-09-12
- DeepSeek V4.1-Flash Runs 502GB Model on a Single RTX 5090 at 5-21 tok/s — AccBalanced · 2026-09-12
- Running 100-200 agents daily: disk space is now the bottleneck, not compute — vincent_koc · 2026-09-12
- Orca releases uncensored MLX weights for DeepSeek V4.1 Flash, cutting refusals by 87-96% — AccBalanced · 2026-09-12