llama.cpp PR Boosts Intel Battlemage Decode Speed by up to 169% at 118K Context
BTA_Labs · reddit · 2026-08-08
A recent llama.cpp PR (#26689) brings massive performance gains for Intel Battlemage GPUs in long-context inference by tweaking a tiny SYCL FlashAttention dispatch logic.
- Core Change: For quantized KV caches (q40 / q80), the decode path is switched from the forced VEC kernel to the TILE kernel.
- Benchmarks: At a 118K context length, decode speeds for Qwen3.6-35B and Gemma 4 12B increased by 127.9% to 168.7%. At 32K context, improvements range from 42% to 74%.
- Caveats: The PR is currently unmerged, benchmarks are mostly author-reported, and it specifically targets quantized KV caches (F16 is unaffected). The community is calling for independent verification on B580/B70 hardware.
More from Infra
- JSON is Burning Your CPU: An Engineering Breakdown of Parse Tax — techNmak · 2026-08-08
- parakeet.wgsl: Transcribes 1 Hour of Audio in 20 Seconds via WebGPU — hamza_q_ · 2026-08-08
- Gary Marcus: The Rise of Neurosymbolic AI Will Bring CPUs Back into the Hardware Mix — Gary Marcus · 2026-08-08
- Fluidstack Hiring: Building Gigawatt-Scale AI Data Centers Like WWII Shipyards — MxMnr · 2026-08-08
- Self-Improving Agents Optimize Inference Stack, Achieving 18% Speedup on B200s — yisongyue · 2026-08-08
- KerasHub Natively Integrates vLLM with Built-in Speculative Decoding — fchollet · 2026-08-08