Running a 139B Model on 6x MI50 GPUs
Old_Grapefruit8774 · reddit · 2026-07-10
The post shares llama-bench results for running MiniMax M2.7 REAP 139B Q3KL on 6x MI50 / Pro VII GPUs, providing throughput data for various scenarios like pp512, tg128, and pp4096+tg128.
The author also lists the hardware specs and launch parameters, including ROCm, llama.cpp-hip, and -ngl 999. This is a detailed log of local inference deployment and performance tuning.
More from Infra
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- Milled from Solid Aluminum: AI Rig Multi-GPU Case for Local Compute — dee_hw · 2026-07-21
- FutureCaribbean’s Buildathon offers $50K, H200 compute, and an NYSE pitch — HeyAmit_ · 2026-07-21
- A new series tests which data-science workflows can run on GPUs today — pandeyparul · 2026-07-21
- Former AWS operator says Bedrock margins can beat SageMaker as agentic AI lifts CPU demand — RihardJarc · 2026-07-21
- Engram shows how agent memory can keep, rewrite, or delete facts asynchronously — philipvollet · 2026-07-21