Qwen3.8-27B crawls at 5 tokens/s on 8GB VRAM + 32GB RAM: best config?

SoAp9035 · reddit · 2026-08-15

A Reddit user asks for help: running Qwen3.8-27B on a consumer rig with 8GB VRAM + 32GB RAM yields only about 5 tokens/s.

They want recommended settings/configs to squeeze out the best possible speed on this hardware, and whether usable speeds are even achievable — a typical local-LLM deployment tuning question.

Original post →

More from Infra

Infra channel →