Why did increasing context size increase speed in Llama.cpp?
satnl · reddit · 2026-09-01
User reports a speed anomaly on RX 9070 XT: running with 131k context is significantly faster (778 t/s) than 65k context (137 t/s). Provides full logs and launch flags, asking if any rule was broken or if this is expected behavior regarding speculative decoding.
More from Infra
- Qwen3.8 Flash hits 415 tok/s on dual DGX Sparks — NVIDIAAI · 2026-09-01
- OpenAI's 'Jalapeno' Chip Revealed: 1500 Tokens/s Throughput — firstadopter · 2026-09-01
- TensorSharp vs llama.cpp: Qwen 3.8 Flash Next Benchmarks — fuzhongkai · 2026-09-01
- Tencent Hunyuan AngelSlim: Compressing Hy4 Model to 214GB with Heterogeneous Inference — 腾讯混元 · 2026-09-01
- Samsung shifts to 8-layer HBM4E for Nvidia with ~20% higher speed spec — 创业邦 · 2026-09-01
- Linux kernel update enables Mac to Linux box connection via USB-C — No-Name-Person111 · 2026-09-01