Four RTX 3080s hit 69 tok/s on Qwen3.6-27B for about $2,000
starkruzr · reddit · 2026-07-23
A benchmarker tested four 20GB RTX 3080s on Vast AI for code generation with Qwen3.6-27B and found the setup surprisingly strong.
- Reported around 69 tokens/sec decode at near-max context (256K) with MTP on
- Prefill dropped to 893 at max context with prompt cache off
- The author says the cards can cost about $400 each, with the full machine roughly $2K all-in
- Conclusion: a cheap 4×3080 box can run dense code models fast, with little quantization and no obvious degradation
More from coding & agent
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- Steal this idea: prompt-to-hardware where agents assemble custom devices — paraschopra · 2026-09-11
- Model Is the Least Interesting Part: A Guide to Six Core AI Architectures from RAG to Multi-Agent — goyalshaliniuk · 2026-09-11
- Non-coder builds layered memory architecture: 20k tokens tracks a year of agent conversations — matteoianni · 2026-09-11
- Warp's six non-engineering teams all run on Linear and Claude Code — mon__lim · 2026-09-11
- 9-year backend dev: AI code isn't the problem, the rate of making a mess is — Sweaty-Landscape-561 · 2026-09-11