Qwen3.8 Flash Next on one RTX 5090 hits 50 t/s decode via FreeToken expert caching
dir3ctly · reddit · 2026-09-20
A Reddit user shares benchmarks running Qwen3.8 Flash Next locally on a single RTX 5090 with FreeToken: 50 t/s token generation and 2300 t/s prefill, stable over long contexts.
- Motivation: llama.cpp still lacks MoE Expert Caching (several unmerged PRs), making default performance disappointing.
- FreeToken's expert caching is the key; the post includes the full launch config (--moe-strategy auto --moe-cache-auto --expert-load parallel, 229376-token sequence override, etc.) using the nvidia/Qwen3.8-Flash-Next-NVFP4 model (convert to FreeToken format or startup is slow).
- Hardware: RTX 5090, 128GB DDR5 6000 (96GB works fine).
- The approach works with other MoE models too; the project is early — vision just landed, no KV quantization or MTP yet.
More from Infra
- On-device model Cactus Needle trends on Hugging Face with tool-calling support — Cactus-Compute · 2026-09-20
- Are data centers dodging taxes? Tax Foundation data on $1B facilities says no — AndyMasley · 2026-09-20
- AMD carries out serious software optimizations for Kimi-K3, analyst says — AccBalanced · 2026-09-20
- RAM price jumps from $350 to $640 in months, pricing out new PCs — chrisalbon · 2026-09-20
- NVIDIA engineer breaks down why DeepSeek re-engineered V4.1 Flash for speed — thursdai_pod · 2026-09-20
- $140 Radeon MI50 paired with GTX-1080Ti boosts local 27B-35B LLM speeds up to 9x — tabletuser_blogspot · 2026-09-20