llama.cpp AMD GFX906 fork: +14% PP, +9% long-context fill vs upstream

milpster · reddit · 2026-09-01

The author released a llama.cpp fork optimized for AMD GFX906 (Radeon VII/MI50/MI60). It achieves +14.1% in first-batch PP and +9.3% in 120k-context fill over the mainline, with tied deep-context TG. The optimizations include reworking the small-Q Flash Attention path and adding adaptive native/convert selection.

Original post →

More from Infra

Infra channel →