Running Qwen3.8-Flash on RTX 3090: Quantization & Performance
crusaderky · reddit · 2026-08-28
A detailed report on running Qwen3.8-Flash-next on an RTX 3090 with 12GB VRAM. The author achieved 160 tok/s prefill and 16 tok/s decode using IQ4XS weights and kvarn5 KV cache. The post includes variant setups for 16GB/12GB cards and links to specific llama.cpp builds and deployment recipes.
Related event: Users Share Qwen3.8 Flash Next Local Deployment Results(2 posts)→
More from Infra
- Essay: Data centers are next-generation ports — communities will build around the flow of intelligence — McDonaghMatthew · 2026-08-29
- Video explainer: how data centers actually drive electricity prices lower — aronchick · 2026-08-29
- Polymarket prices 68% chance a US state enacts a data center moratorium in 2026 — Polymarket · 2026-08-29
- Microsoft reportedly reassures employees on AI data center energy and emissions impact — Polymarket · 2026-08-29
- Google Cloud SQL Introduces Performance Assessments Preview — rseroter · 2026-08-29
- Decathlon Switches to Chronos-2 for Large-Scale Demand Forecasting on AWS — AWS ML Blog · 2026-08-29