Is 5 tokens/s usable for local LLMs? Redditor runs 27B model off an iGPU

Zombiecidialfreak · reddit · 2026-09-10

A Redditor shares two ways of running Qwen3 27B locally: a split across a 3060 12GB and 9070XT at 20t/s, or entirely off a 780m iGPU with 5400MHz DDR5 at just 5t/s. He finds the slower setup more practical since the system stays fully usable for gaming and daily tasks, trading 4x longer responses for a free machine — sparking debate on what token speed counts as usable for local inference.

Original post →

More from Infra

Infra channel →