Running Qwen3.8-Flash on RTX 3090: Quantization & Performance

crusaderky · reddit · 2026-08-28

A detailed report on running Qwen3.8-Flash-next on an RTX 3090 with 12GB VRAM. The author achieved 160 tok/s prefill and 16 tok/s decode using IQ4XS weights and kvarn5 KV cache. The post includes variant setups for 16GB/12GB cards and links to specific llama.cpp builds and deployment recipes.

Related event: Users Share Qwen3.8 Flash Next Local Deployment Results(2 posts)→

Original post →

More from Infra

Infra channel →