New llama.cpp PR assigns four GDN state columns per warp for another Qwen 3.x prefill speedup

jacek2023 · reddit · 2026-10-08

Developer SongXiaoXi opened PR #30087 on llama.cpp optimizing prompt processing by assigning four GDN state columns per warp, marking another speedup for Qwen 3.x models. The author quips that soon your Qwen will read your entire project before you can blink.

Original post →

More from Infra

Infra channel →