Local LLM inference speeds jump ~10x in a week: single consumer GPU now hits 2200 prefill

oran_ge · x · 2026-10-03

Quoting @ivanalogcom's hands-on report: a week ago, dual RTX 5070 Ti cards locally running Qwen Flash managed only 200 prefill/10 decode tokens. Today, with the Strata architecture and a custom PR, a single card reaches 2200/67 — a near-10x speedup on both read and write, a level previously reserved for RTX 5090, DGX Spark, or M3 Ultra rigs, and beyond most model APIs.

The author argues this is a huge win for local inference, possibly a converged optimum for current hardware — and a potential death sentence for businesses selling mediocre model compute via API.

Original post →

More from Infra

Infra channel →