Qwen3.8 Flash Next hits 49 tok/s locally on 2x RTX 3090 with FlashNext llama.cpp fork

whiteh4cker · reddit · 2026-09-11

Reddit user whiteh4cker benchmarked Qwen3.8 Flash Next UD-Q4KXL locally on 2x RTX 3090 (Windows 11, 192GB DDR5) using the FlashNext fork of llama.cpp, boosting generation speed from 20 t/s on the main branch to 49 t/s, with 140 t/s prompt processing (faster on main branch).

The post includes full repro details:

A ready-to-copy config for anyone squeezing multi-GPU local inference out of the newest MoE models.

Original post →

More from Infra

Infra channel →