Local LLM inference speeds jump ~10x in a week: single consumer GPU now hits 2200 prefill
oran_ge · x · 2026-10-03
Quoting @ivanalogcom's hands-on report: a week ago, dual RTX 5070 Ti cards locally running Qwen Flash managed only 200 prefill/10 decode tokens. Today, with the Strata architecture and a custom PR, a single card reaches 2200/67 — a near-10x speedup on both read and write, a level previously reserved for RTX 5090, DGX Spark, or M3 Ultra rigs, and beyond most model APIs.
- RTX 4090/5090 also see big gains; 5090 with streaming is approaching RTX 6000 speeds (VRAM remains the limit)
- The surge is driving consumer hardware price hikes — even the obscure Intel B70 32G card has sold out
- The gains come from DeepSeek/Qwen Flash's design (splitting intelligence into attention, experts, and lookup) plus Strata's pipelined architecture that stages different capabilities across storage tiers
The author argues this is a huge win for local inference, possibly a converged optimum for current hardware — and a potential death sentence for businesses selling mediocre model compute via API.
More from Infra
- SemiAnalysis asks if GPUs are money printers: TPU v7, Vera Rubin vs Blackwell — AccBalanced · 2026-10-03
- gufo Ships Pre-packaged Windows Build for AMD Strix Halo, 40 tps on Agentic Workloads — hiImMate · 2026-10-03
- Musk on compute scaling: sufficient quantity is a quality all its own — XFreeze · 2026-10-03
- Can a RTX 5080 gaming PC run a local coding agent comparable to Codex? — evilgu · 2026-10-03
- Inference now bigger than training, and it's reshaping data center buildouts — AccBalanced · 2026-10-03
- Terafab reportedly targets 1 TW of compute per year, ~50x current global AI output — XFreeze · 2026-10-03