llama.cpp PR adds missing AMD GCN MMQ config, boosting MI50/MI60 inference
pmttyji · reddit · 2026-09-12
A llama.cpp pull request (#27841) adds the missing AMD GCN MMQ configuration for the ROCm/HIP backend.
According to the contributor, the change delivers prompt processing (PP) improvements for RDNA2 cards like the MI50 and MI60, with updated pp t/s benchmarks posted in the bottom comments.
More from Infra
- Best local uncensored vision+text models for an RTX 5070 Ti 16GB setup? — Mystvearn_ · 2026-09-12
- From Accuracy to Latency: Why Inference Engineering Is a Different Discipline — mdancho84 · 2026-09-12
- CheckCle: self-hosted open-source full-stack monitoring platform hits 2.9k GitHub stars — tom_doerr · 2026-09-12
- AI in Space Is Mostly Inference: '$10/Month 200-IQ Employees' Means Infinite Demand — JOBhakdi · 2026-09-12
- Building a Coding Agent? Modal Sandboxes Shine, but E2B and Vercel Remain Unproven — pauliusztin · 2026-09-12
- Same batch job costs $97 on Claude Sonnet vs $13 on a rented H200 — 7-8x cheaper — pauliusztin · 2026-09-12