AI inference should report prefill, decode, and KV cache separately

isidentical · x · 2026-07-29

The post argues that inference performance should be reported with three separate metrics: prefill throughput, decode throughput, and KV cache size, all measured at a given context length.

It criticizes a single aggregate tokens-per-second number as too variable to compare meaningfully and closer to a marketing metric than a useful benchmark.

Original post →

More from Infra

Infra channel →