Running Qwen3 27B locally on two 24GB machines: 3-4 tok/s and tasks failing midway
shadyshak · reddit · 2026-09-16
A Reddit user benchmarked Qwen3 8:27B via Ollama on two 24GB machines (Minisforum UM790 Pro and a Ryzen 9 7940HS with Radeon 780M) with the same coding prompt.
Results: eval speeds of only 3.2 and 4.7 tokens/s, full runs taking 11-13 minutes, and — worse — neither instance completed the task. One stopped mid-generation and produced nothing on "continue"; the other's output got truncated to almost no tokens. The two instances also diverged in strategy (one listing options, one coding directly).
A cautionary data point for anyone attempting 27B-class local models on small-memory hardware: expect slow generation and incomplete outputs.
More from Infra
- Earendil launches Radius, an inference layer bringing tokens, routing and search to Pi — HankYeomans · 2026-09-16
- Multi-turn agent RL training at scale on HF Hub: 9,523 sandboxes in 14h, zero crashes — vanstriendaniel · 2026-09-16
- Apple Weighs M8-Based AI Server With Nvidia NVLink to Reenter Server Market — gappyvalley · 2026-09-16
- uv 0.12.9 reworks ZIP extraction to speed up cold-cache Python installs — KhuyenTran16 · 2026-09-16
- Apple reportedly building enterprise AI server on its own chips, launching 2029 — Polymarket · 2026-09-16
- Citi analysts: HBM demand to surge 62% in 2027 and another 69% in 2028 — TiernanRayTech · 2026-09-16