gfx906-llama-cpp Fork Boosts MI50/MI60 Inference +23% Prefill, Fits 250k Context on 40GB
milpster · reddit · 2026-09-05
Developer milpster shipped a major update to gfx906-llama-cpp, a fork optimizing llama.cpp for AMD GCN cards like MI50/MI60/Radeon VII: prefill PP16384 up from 332.5 to 410 t/s (+23%), 120k deep fill +14%, TG decode +11% (15.1 t/s), and 250k context squeezed into 40GB VRAM (upstream can't fit it). Outputs are bit-identical to upstream. Gains came mostly from adopting relevant upstream PRs; the author notes the fork was built with AI assistance. A practical resource for cheap local LLM inference on aging GPUs.
More from Infra
- Dev Tests NVIDIA's Deepseek V4 NVFP4: 1M Context, 96% Memory at Batch 2048 — HankYeomans · 2026-09-05
- Open-Source On-Device Face Swap Hits Android: Real-Time on Hexagon NPU, 66 MB APK — Few_Caregiver8134 · 2026-09-05
- Who Needs a GB300? Researcher Makes a Movie on a $500 16GB GPU — francoisfleuret · 2026-09-05
- Minima quantizes all 496 layers of Qwen3.8-27B to NVFP4 W4A4, matching BF16 at 2.9x smaller — pbaylies · 2026-09-05
- 19 latency patterns to cut non-model latency in AI applications — bibryam · 2026-09-05
- Astra reportedly trained on 100,000 GPUs, a staggering compute scale — sudoraohacker · 2026-09-05