Running Llama Locally: ~2-8 tokens/s on CPU, Boosted by GPU

thisdudelikesAI · x · 2026-07-07

Running quantized LLMs purely on a CPU yields about 2-8 tokens/second—functional but relatively slow. Performance sees a significant boost when hardware acceleration is added via GPUs (NVIDIA, AMD, or Apple Silicon's unified memory architecture).

Original post →

More from Infra

Infra channel →