antirez: stop obsessing over t/s — short thinking phases and fast prefill are what matter
antirez · x · 2026-09-29
Redis creator antirez argues generation tokens-per-second is the wrong headline metric for local inference. He predicts Chinese open model providers will realize the real game is minimizing the thinking phase — with fast prefill, 20/30 t/s generation is plenty.
Related event: antirez: Prefill Speed Matters More Than Generation Speed for Local LLMs(2 posts)→
More from Infra
- Chinese CNC factory declines part over export controls; poster builds it anyway — wavefnx · 2026-09-29
- Ben Lorica: the AI data problem has moved downstream — from raw material to usable-data labor and licensed access — bigdata · 2026-09-29
- Researchers Build Backdoor Method to Rank AI Model Energy Use as Big Three Stay Silent — relianceschool · 2026-09-29
- Agent spins up 322 Hugging Face Jobs in 90 minutes for about $4 — victormustar · 2026-09-29
- 80,000+ proxy servers hide stolen AI credentials fueling LLM abuse, Team Cymru finds — ChuckDBrooks · 2026-09-29
- Investor Calls for a 'Great Wall of Compute' to Export Intelligence, Not Dollars — McDonaghMatthew · 2026-09-29