llama.cpp adds Vulkan INT8 coopmat matmul for AMD RDNA3/4, boosting 7900XTX prefill 25%+
nickm_27 · reddit · 2026-09-24
A merged llama.cpp change adds a Vulkan int8 coopmat1 matmul implementation for AMD RDNA3/RDNA4 GPUs. Benchmarked on a 7900XTX with gemma4 26B.A4B Q40, pp512 jumped from 3410.5 to 4331.2 t/s (27% faster), and long-context prefill at d16384 rose from 2130 to 2402 t/s; tg128 also improved slightly. A free, meaningful speedup for local inference on AMD cards.
More from Infra
- Xiaomi's MiMo-V2.6 Pro tops open-source leaderboard at 46 AA, $0.13 per task — SucceededMind · 2026-09-24
- Celesto Launches Real Computers for AI Agents With 500ms-Boot MicroVMs — aniketmaurya · 2026-09-24
- Report: Google to Launch Space TPUs on Falcon 9 Next Week to Test Orbital AI Data Centers — ns123abc · 2026-09-24
- Researchers exploited Cloudflare Containers flaw to read other tenants' residual disk data — matthew_d_green · 2026-09-24
- IIT Delhi Develops India's First Indigenous Micro GPU — Paimaamu · 2026-09-24
- Marvell and GlobalFoundries sign multi-year deal to expand silicon germanium capacity — Beth_Kindig · 2026-09-24