New llama.cpp PR assigns four GDN state columns per warp for another Qwen 3.x prefill speedup
jacek2023 · reddit · 2026-10-08
Developer SongXiaoXi opened PR #30087 on llama.cpp optimizing prompt processing by assigning four GDN state columns per warp, marking another speedup for Qwen 3.x models. The author quips that soon your Qwen will read your entire project before you can blink.
More from Infra
- After a $10,811 Cloudflare bill, dev ditches delayed billing alerts for a Grok-built 3-hour usage monitor — mazzaTalk · 2026-10-08
- Running an LLM town with 800+ persistent agents: concurrency, caching, and costs — Low_Bad_6585 · 2026-10-08
- AI borrowing costs hit ~11% as JPMorgan markets $5B Volta loan for 36,000 Nvidia GPUs — mjdramstead · 2026-10-08
- GLM 5.3 Flash Served on 2 DGX Sparks: Open Recipe Hits 77.6 tok/s with 3 — EAccelerate_42 · 2026-10-08
- tinygrad launches new tinybox deep learning rig, configurable up to 4 GPUs — GiorgioPatrini · 2026-10-08
- Inference Overtakes Training as Biggest AI Market, Reshaping Data Center Buildouts — FinanceYF5 · 2026-10-08