antirez: stop obsessing over t/s — short thinking phases and fast prefill are what matter

antirez · x · 2026-09-29

Redis creator antirez argues generation tokens-per-second is the wrong headline metric for local inference. He predicts Chinese open model providers will realize the real game is minimizing the thinking phase — with fast prefill, 20/30 t/s generation is plenty.

Related event: antirez: Prefill Speed Matters More Than Generation Speed for Local LLMs(2 posts)→

Original post →

More from Infra

Infra channel →