Running LLMs on mixed AMD GPUs: 11 t/s inference speed achieved
Hyungsun · reddit · 2026-08-02
A developer on Reddit shared benchmark results running a quantized version of DeepSeek-V4-Flash-0731 on a mixed AMD GPU setup (1x Radeon 7900 XTX 24GB + 3x Instinct MI60 32GB) with 128GB DDR4 RAM.
- Hardware: Dual Intel Xeon E5-2650 v4 CPUs. The heterogenous GPU combo was a result of one MI60 failing.
- Performance: Using an unoptimized llama.cpp (ROCm backend), prompt processing hit 140 t/s (dropping to 85 t/s at 64k context), with generation speeds steady at around 11.9 t/s.
More from Infra
- AMD MI355X Beats NVIDIA B200 in Kimi K3 Deployment with 952 tok/s — adrianscottcom · 2026-08-03
- MiniMax H3 Gets Day 0 Support in SGLang, Runs Locally on Dual RTX 5090s — ying11231 · 2026-08-03
- Qwen3.8-27B Open Weights Coming, Runs Locally on 17GB RAM — danielhanchen · 2026-08-03
- AirLLM Breaks VRAM Barrier: Runs 70B LLMs on a Single 4GB GPU — techNmak · 2026-08-03
- MiniMax H3 Open Weights Hit fal with Out-of-the-Box Inference Optimizations — gorkem · 2026-08-03
- ComfyUI Adds Day 0 Support for MiniMax Video Model, Slashing VRAM by 66% for RTX 3060 — crystal_alpine · 2026-08-03