llama.cpp Updates: Mamba-2 Prefill Speedup Over 20%, Metal Backend Optimized
pmttyji · reddit · 2026-07-28
llama.cpp has recently merged several key performance optimizations and fixes:
- Mamba-2 Prefill Acceleration: Added chunked SSD matmul (#22675), boosting long-context (e.g., 8k) prefill speeds by over 20% on Nemotron-Nano-9B-v2.
- Apple Metal Optimization: Introduced FWHT kernel for the Metal backend (#25924), increasing low-precision (e.g., q80) generation speeds by around 3.5% for DeepSeek-V4-Flash on M4 Max.
- Speculative Decoding: Added Eagle3-v3 support for GPT-OSS models (#25794).
- Bug Fixes: Resolved an SDPA scale issue in the SYCL backend and restored iq4nl support for Vulkan.
More from Infra
- Screenshot shows Anthropic crawl spikes as users speculate Sonnet 6 training is underway — marclou · 2026-07-28
- Moonshot’s Kimi K3 goes live on Modal with $3 in, $15 out pricing — ivan_bezdomny · 2026-07-28
- NVIDIA says Jetson now fits in a bag while powering robots and edge AI — nordicinst · 2026-07-28
- India can assemble a finished chip in Gujarat, but the core is still imported — santoshpanda · 2026-07-28
- Open competition targets 36.8% faster Laguna XS 2.1 inference on consumer Macs — gajesh · 2026-07-28
- Make electricity abundant first, then AI abundance will follow — XFreeze · 2026-07-28