Running Qwen3.8 Locally on 6x3090s with exllamav3 Hits 80-120 tok/s

takoulseum · reddit · 2026-09-26

A redditor shares a local-agent setup: Qwen3.8 in 6bpw exl3 quant (by turboderp) served via the exllamav3 engine on 6x RTX 3090s delivers roughly 80-120 tok/s with good prefill and quality — "crazy for CUDA." They control the Hermes agent daily from their phone over Matrix, and use it beyond delegating to coding harnesses like opencode, assuming lower quants on fewer GPUs would work similarly.

Original post →

More from coding & agent

coding & agent channel →