antirez: stop chasing tokens/s — fast prefill and short thinking beat 20-30 t/s myths
antirez · x · 2026-09-29
antirez argues that tokens per second is the wrong primary metric for local LLM inference: Chinese open-model providers will soon compete on making the thinking phase as short as possible, and 20/30 t/s is enough with fast prefill. He adds that upcoming models will make DGX Spark and Strix Halo far more compelling, and that local hardware shouldn't be evaluated in a vacuum — its value depends on the class and strengths of models being released.
Related event: antirez: Prefill Speed Matters More Than Generation Speed for Local LLMs(2 posts)→
More from Infra
- OpenAI Engineer Details Six Prompt Caching Improvements Cutting Latency and Cost — pvncher · 2026-09-29
- InstaCloud launches agent-native serverless cloud that lets Claude Code and Cursor provision infra via one command — testingcatalog · 2026-09-29
- Enrichment bottlenecks nuclear, nuclear bottlenecks power, power bottlenecks AI — le_james94 · 2026-09-29
- Turso 0.8.0 lands: up to 7x faster than SQLite with 500x lower tail latency — glcst · 2026-09-29
- AMD ships Ryzen AI Max+ PRO with 192GB unified memory, 50% more than NVIDIA's upcoming Spark — ryanshrout · 2026-09-29
- Founders now spend their time hunting compute: Instinct demand doubles every week — ai · 2026-09-29