Why did increasing context size increase speed in Llama.cpp?

satnl · reddit · 2026-09-01

User reports a speed anomaly on RX 9070 XT: running with 131k context is significantly faster (778 t/s) than 65k context (137 t/s). Provides full logs and launch flags, asking if any rule was broken or if this is expected behavior regarding speculative decoding.

Original post →

More from Infra

Infra channel →