Qwen 3.8 Next Flash at 3.05bpw EXL3 runs like Q8 on 3x RTX 3090s, dev reports
nicholas_the_furious · reddit · 2026-09-20
A local LLM enthusiast shares detailed benchmarks running Qwen 3.8 models on 3x RTX 3090s (250-265W, one on a TB4 eGPU):
- Daily driver: Qwen 3.8 27B UD-Q8KL on 2x 3090s in TP, FP16 KV cache, 250k+ context with no perceived degradation (vs 150k for 3.6), great as a coding agent.
- Qwen 3.8 Flash Next UD-Q4KXL produced a stunning Flappy Bird variant with full UI, animations, and moving pipes — far more world knowledge than the 27B — but RAM offload killed speed (35-50 t/s out, low-hundreds PP).
- The fix: EXL3 3.05bpw quant fits fully in VRAM with full context and FP16 KV cache, feels indistinguishable from Q8, and delivers 110+ t/s output and 1500+ t/s prefill.
Verdict: if you have 3x 3090s, run Qwen 3.8 Next Flash at exl3 3.05bpw; EXL3 isn't worth it for the 27B, and Unsloth Q3 quants may be worth comparing.
More from Infra
- Rumor: Anthropic trails OpenAI in training compute intensity and inference economics — teortaxesTex · 2026-09-20
- Running Qwen 3.8 Next on six V100s: MTP nearly doubles output to 43 tok/s — Odd_Caterpillar_2994 · 2026-09-20
- Ben Bajarin: Agentic AI will spawn an 'agentic native' CPU tier in datacenters — BenBajarin · 2026-09-20
- Redditor pleads with FP4 inference engine builders: small dense models at FP4 are cooked — buttplugs4life4me · 2026-09-20
- Chips will depreciate more slowly as materials change, and thermal compute is still coming — beffjezos · 2026-09-20
- Devs debate running stateful AI agent runtimes on Cloudflare Workers and other edge runtimes — merlinofthewater · 2026-09-20