llama.cpp PR tunes flash attention shapes for faster Gemma 26B A4B inference
jacek2023 · reddit · 2026-09-23
PR #28450 in ggml-org/llama.cpp optimizes the flash attention tensor shapes for gemma4-26b-a4b, yielding a notable speedup for local inference.
This kind of low-level shape tuning is a standard community technique for squeezing performance out of llama.cpp, and offers direct gains for users running Gemma-class models on their own hardware.
More from Infra
- Distillation Costed ~$16 Total: It's a Data-Cost Problem, Not a Training-Cost Problem — Gradio · 2026-09-23
- Chutes names new CEO, pretrains 8B MoE in public, hits 15.5k tok/s on one RTX 5090 — markjeffrey · 2026-09-23
- Halo post-training framework claims 2.8x TRL throughput, accused of cherry-picking benchmarks — _ScottCondron · 2026-09-23
- OpenAI credits caching and inference gains for GPT-6's 50% price cut — OpenAIDevs · 2026-09-23
- How OpenAI Built GPT-Live: Engineers Deep-Dive with ByteByteGo — juberti · 2026-09-23
- China data center power to hit 774 TWh by 2030, US 649 TWh, forecasts say — teortaxesTex · 2026-09-23