llama.cpp PR fuses Qwen activation chain for faster Vulkan inference
jacek2023 · reddit · 2026-09-28
A llama.cpp pull request (#29520 by fxgsell) fuses the SCALE -> SIGMOID -> SCALE -> hcpost activation chain for qwen4exp on the Vulkan backend, cutting overhead. The author calls it the next Qwen speedup for Vulkan users running local deployments.
More from Infra
- MLX MoE Layer Gets 1.5x Faster via Better Tile Scheduling in Grouped Matmul — awnihannun · 2026-09-28
- Cloudflare incident: skipped block zeroing leaked tenant data across 18 of 24 containers — arpit_bhayani · 2026-09-28
- On RTX 5090, Qwen 27B hits 200 TPS but Flash next only 50: what model sits between for coding? — MasterNomie · 2026-09-28
- Lumen Launches On-Demand Dedicated Internet Up to 100 Gbps at 10M US Sites — shashib · 2026-09-28
- Local AI comes in two flavors: laptop-scale for the masses vs SMB on-prem setups — TheZachMueller · 2026-09-28
- Cloudflare's agent-first Kitesurf browser adds WebMCP, passes 730k WPT subtests, runs in terminals — Cloudflare Blog · 2026-09-28