Four RTX 3080s hit 69 tok/s on Qwen3.6-27B for about $2,000
starkruzr · reddit · 2026-07-23
A benchmarker tested four 20GB RTX 3080s on Vast AI for code generation with Qwen3.6-27B and found the setup surprisingly strong.
- Reported around 69 tokens/sec decode at near-max context (256K) with MTP on
- Prefill dropped to 893 at max context with prompt cache off
- The author says the cards can cost about $400 each, with the full machine roughly $2K all-in
- Conclusion: a cheap 4×3080 box can run dense code models fast, with little quantization and no obvious degradation
More from coding & agent
- TerraLingua opens public access to autonomous agents in a persistent world — kenneth0stanley · 2026-07-23
- An Ordinals marketplace was built entirely with an AI tool — adamamcbride · 2026-07-23
- A practical eval loop for coding agents starts with codebase traces and user feedback — EdenEmarco177 · 2026-07-23
- Sierra’s internal agent hit the same bottleneck: getting the right context — blaizedsouza · 2026-07-23
- An AI skill only stays if it works daily, survives model switches, and stays cheap to maintain — Numerous-Service-980 · 2026-07-23
- GPT-5.6 Sol and Fable 5 test car-building in Scrap Mechanic with an MCP — Angaisb_ · 2026-07-23