Running Qwen3 27B locally on two 24GB machines: 3-4 tok/s and tasks failing midway

shadyshak · reddit · 2026-09-16

A Reddit user benchmarked Qwen3 8:27B via Ollama on two 24GB machines (Minisforum UM790 Pro and a Ryzen 9 7940HS with Radeon 780M) with the same coding prompt.

Results: eval speeds of only 3.2 and 4.7 tokens/s, full runs taking 11-13 minutes, and — worse — neither instance completed the task. One stopped mid-generation and produced nothing on "continue"; the other's output got truncated to almost no tokens. The two instances also diverged in strategy (one listing options, one coding directly).

A cautionary data point for anyone attempting 27B-class local models on small-memory hardware: expect slow generation and incomplete outputs.

Original post →

More from Infra

Infra channel →