Is 5 tokens/s usable for local LLMs? Redditor runs 27B model off an iGPU
Zombiecidialfreak · reddit · 2026-09-10
A Redditor shares two ways of running Qwen3 27B locally: a split across a 3060 12GB and 9070XT at 20t/s, or entirely off a 780m iGPU with 5400MHz DDR5 at just 5t/s. He finds the slower setup more practical since the system stays fully usable for gaming and daily tasks, trading 4x longer responses for a free machine — sparking debate on what token speed counts as usable for local inference.
More from Infra
- Keras ships ZeroModels: 100+ model families in pure Keras 3, runnable on any backend — fchollet · 2026-09-10
- Cohere moves to NVIDIA Blackwell, cutting token costs and TTFT by 30–50% — cohere · 2026-09-10
- Author uses local AI models to review book manuscripts for $0 in tokens — walkingriver · 2026-09-10
- Musk's Colossus data center fuels massive local backlash, reports Scientific American — scientificamerican · 2026-09-10
- Analyst: Huawei to undercut US AI stack with cheaper chips tuned for Chinese models — 2C_ornot2C · 2026-09-10
- Photon 2.2 Ships Optimized Local Inference for Ampere Through Blackwell GPUs — JFPuget · 2026-09-10