Longer Context = Faster Prefill? A Puzzling llama.cpp Benchmark Anomaly
Ekepa · reddit · 2026-09-11
A local-deployment user observed on llama.cpp (RX 6600 XT 8GB, Vulkan, all 43 layers offloaded) that prefill speed is not monotonically decreasing with context length: after 935→318 tok/s from 1018–8151, throughput rebounds — 16302 context is faster than 8151 (472 vs 318 tok/s), and 130416 context (89.7 tok/s) beats 65208 (76.7 tok/s) despite processing twice the tokens. Run-to-run spread is under 2%, so the anomaly reproduces consistently. The cause is unknown; the author is asking the community for explanations.
More from Infra
- Training a 6-Expert MoE GPT-2 From Scratch on a Single RTX 3090 in 8 Days — rasbt · 2026-09-11
- B3IQ Sells Eight Figures of GPUs in Two Weeks, Bets AI Infra Is a $100B Market — templecrash · 2026-09-11
- It Cost $100 in API Credits for an AI Agent to Install Free Software — MartinGTobias · 2026-09-11
- bartowski unveils per-tensor layout maps for GGUF quantization, tests show across-the-board gains — noneabove1182 · 2026-09-11
- antirez runs DeepSeek v4.1 Flash locally on a 128GB M5 Max, SSD streaming surprisingly fast — antirez · 2026-09-11
- OpenAI could 7x its training compute tomorrow: why open-source models still trail by one generation — soumitrashukla9 · 2026-09-11