Qwen3.8-Flash-Next only hits 15 tok/s on 4x RTX 5060 Ti 16GB setup
Ambitious_Fold_2874 · reddit · 2026-09-12
- Running Qwen3.8-Flash-Next Q80 GGUF on 4x RTX 5060 Ti 16GB (quad-channel DDR4) via Unsloth Studio yields only 15 t/s generation and 90 t/s prompt processing — well below expectations.
- llama.cpp errored out; the poster shares a full llama-server command using the MTP draft model for speculative decoding (--spec-type draft-mtp, max 3 drafts), q80 KV cache, 256K context.
- Quad-channel DDR4 memory bandwidth is likely the bottleneck — a useful data point for consumer multi-GPU local deployments.
More from Infra
- Musk announces Terafab: Tesla, SpaceX and xAI to build 1TW/year chip fab — elonmusk · 2026-09-12
- The hidden cost of agents is KV cache: DeepSeek compresses to ~890 bytes per token — altryne · 2026-09-12
- A ~$3,157 dual RTX 3090 inference rig hits 70 tps on Qwen3 27B — Puzzleheaded_Ad_8575 · 2026-09-12
- AMD ships DeepSeek v4.1 Flash support 2 days late, up to 42x worse perf per dollar vs B200 — IanAndrewsDC · 2026-09-12
- Deep dive: OpenAI's Jalapeno inference accelerator architecture from Hot Chips 2026 — bookwormengr · 2026-09-12
- h3 studio: native Metal web UI for MiniMax-H3 video gen on Apple Silicon, model stays resident — janishar · 2026-09-12