Qwen3.8 Flash Next IQ1_M hits 55 tok/s on a 5060 Ti 16GB and still codes well
bobaburger · reddit · 2026-10-07
- On a modest 5060 Ti 16GB + 32GB RAM, qwen3.8-flash-next-coder-iq1m with nctx=65k averages 1.5k tok/s prompt processing and 55 tok/s generation
- Despite the extreme IQ1M quant, one-shot quality surprised: a landing page in 2 min at 46 tok/s looked better than Q3-class local models and avoided the typical "Claude-ish" style issues
- An interactive 3D globe took 8 min for v1 plus 1 min to fix JS errors, still impressive
- Live demos: bakery and earth pages on Vercel
More from Infra
- One RTX Pro 6K Sidecar Lifts DGX Station GB300 to 60K tok/s Prefill with DeepSeek V4.1 Flash — Sentdex · 2026-10-07
- Hugging Face's Talk: A Full Tour of llama.cpp and the Local AI Inference Ecosystem — unofficialmerve · 2026-10-07
- Constellation Energy stock surges 14% on 20-year nuclear power deal with Google — Polymarket · 2026-10-07
- DNS root KSK rollover hits October 11: Cloudflare explains KSK-2024 switch and readiness test — Cloudflare Blog · 2026-10-07
- Qwen 3.8 27B hits 96 t/s decode with 110k context on a single 16GB RTX 5080 via NInfer — Kernoriordan · 2026-10-07
- OpenAI to fund Nicholas Nethercote's work speeding up the Rust compiler — charliermarsh · 2026-10-07