Memory Bandwidth Caps LLM Decoding: A Simple Token/s Formula

A Reddit-derived rule of thumb estimates local LLM decoding speed: tokens/s ≈ memory bandwidth / 2N, since decoding is bandwidth-bound rather than compute-bound for dense models.

2026-09-03 ~ 2026-09-05 · 2 related posts