llama.cpp AMD GFX906 fork: +14% PP, +9% long-context fill vs upstream
milpster · reddit · 2026-09-01
The author released a llama.cpp fork optimized for AMD GFX906 (Radeon VII/MI50/MI60). It achieves +14.1% in first-batch PP and +9.3% in 120k-context fill over the mainline, with tied deep-context TG. The optimizations include reworking the small-Q Flash Attention path and adding adaptive native/convert selection.
More from Infra
- Tencent's Hy4 Preview: 1.25-bit Quantization Cuts Model Size by 7x — ccerrato147 · 2026-09-01
- Opinion: Humanoid robot sales predictions are meant to be forgotten — TiernanRayTech · 2026-09-01
- Hugging Face Transformers Adds Activation Checkpoint Offload to Reduce GPU Memory — StasBekman · 2026-09-01
- Meta Open Sources MetaRoCE Protocol for AI-Scale Ethernet Clusters — Meta_Engineers · 2026-09-01
- Amazon quietly hikes prices on Echo, Kindle, citing rising memory costs — film_girl · 2026-09-01
- Cloudflare turns global network into agentic cloud with Dynamic Workers and AI Gateway — dscape · 2026-09-01