Why does nobody benchmark prefill? Local rigs ignore input processing speed vs API providers

BobtheGodGamer · reddit · 2026-09-19

The author points out a blind spot in local LLM benchmarking: when people with Spark or Strix Halo rigs compare against API providers (Opus, Astra, etc.), they focus almost entirely on decode speed while ignoring prefill (input processing) speed — which matters a lot with long contexts if your local prompt processing is slow.

He asks whether any public benchmarks or figures exist comparing how fast API providers process input versus local hardware. A question post, but it highlights a real gap in how local inference performance is usually evaluated.

Original post →

More from Infra

Infra channel →