Longer Context = Faster Prefill? A Puzzling llama.cpp Benchmark Anomaly

Ekepa · reddit · 2026-09-11

A local-deployment user observed on llama.cpp (RX 6600 XT 8GB, Vulkan, all 43 layers offloaded) that prefill speed is not monotonically decreasing with context length: after 935→318 tok/s from 1018–8151, throughput rebounds — 16302 context is faster than 8151 (472 vs 318 tok/s), and 130416 context (89.7 tok/s) beats 65208 (76.7 tok/s) despite processing twice the tokens. Run-to-run spread is under 2%, so the anomaly reproduces consistently. The cause is unknown; the author is asking the community for explanations.

Original post →

More from Infra

Infra channel →