Running a 139B Model on 6x MI50 GPUs

Old_Grapefruit8774 · reddit · 2026-07-10

The post shares llama-bench results for running MiniMax M2.7 REAP 139B Q3KL on 6x MI50 / Pro VII GPUs, providing throughput data for various scenarios like pp512, tg128, and pp4096+tg128.

The author also lists the hardware specs and launch parameters, including ROCm, llama.cpp-hip, and -ngl 999. This is a detailed log of local inference deployment and performance tuning.

Original post →

More from Infra

Infra channel →