antirez: stop chasing tokens/s — fast prefill and short thinking beat 20-30 t/s myths

antirez · x · 2026-09-29

antirez argues that tokens per second is the wrong primary metric for local LLM inference: Chinese open-model providers will soon compete on making the thinking phase as short as possible, and 20/30 t/s is enough with fast prefill. He adds that upcoming models will make DGX Spark and Strix Halo far more compelling, and that local hardware shouldn't be evaluated in a vacuum — its value depends on the class and strengths of models being released.

Related event: antirez: Prefill Speed Matters More Than Generation Speed for Local LLMs(2 posts)→

Original post →

More from Infra

Infra channel →