Qwen3.8-27B crawls at 5 tokens/s on 8GB VRAM + 32GB RAM: best config?
SoAp9035 · reddit · 2026-08-15
A Reddit user asks for help: running Qwen3.8-27B on a consumer rig with 8GB VRAM + 32GB RAM yields only about 5 tokens/s.
They want recommended settings/configs to squeeze out the best possible speed on this hardware, and whether usable speeds are even achievable — a typical local-LLM deployment tuning question.
More from Infra
- RTX 3090 gets 35 t/s on Qwen 3.8 27B — cviperr33 · 2026-08-15
- CME to launch futures contracts tracking Nvidia H100/B100 compute costs — AccBalanced · 2026-08-15
- Feedback: Serverless GPU capacity tight, placement times high — tobowers · 2026-08-15
- Touchmark launches futures marketplace for AI tokens, up to 30% below on-demand rates — ycombinator · 2026-08-15
- Macro Analysis: Is the AI Dip a Buy? Key Levels to Watch — Beth_Kindig · 2026-08-15
- Optimized Dual 3090 Quantization of Qwen3.8-27B Released — luedtek · 2026-08-15