Qwen3.8-Flash on 4x AMD V620 hits 1,300 PP and 70+ tok/s on coding
Thin_Pollution8843 · reddit · 2026-09-17
A Reddit user benchmarked the quantized Qwen3.8-Flash-Next (W4A16-AutoRound, roughly between q5-xl and q6-xl quality) on 4 RDNA2 AMD V620 GPUs via the community vllm-rdna project: 1,300-1,393 tok/s prompt processing and 56-59 tok/s generation at 32K-128K contexts, rising to 68-72.5 tok/s on coding suites. A budget path to local inference on aging datacenter GPUs.
More from Infra
- Cloudflare Launches First Stealth Model Union Alpha, Blending Multiple LLMs Per Request — ritakozlov · 2026-09-17
- Qwen3.8-Flash-Next on SGLang NVFP4: 254K-Context TTFT Drops 35s→22s, Full Gauntlet Tested — FantasticNature7590 · 2026-09-17
- Privacy-focused LLM service Venice hits 250B daily tokens, up 2.5x in months — 0xAllen_ · 2026-09-17
- Baseten launches Hosted Tools, bringing server-side web search to open models — baseten · 2026-09-17
- Dev proposes predictive dynamic context caching for Claude Code — Sauers_ · 2026-09-17
- Starlink expands across Latin America: 10,000 antennas to connect 8,000 schools in Honduras alone — NicoVerderosa · 2026-09-17