Running LLMs on mixed AMD GPUs: 11 t/s inference speed achieved

Hyungsun · reddit · 2026-08-02

A developer on Reddit shared benchmark results running a quantized version of DeepSeek-V4-Flash-0731 on a mixed AMD GPU setup (1x Radeon 7900 XTX 24GB + 3x Instinct MI60 32GB) with 128GB DDR4 RAM.

Related event: DeepSeek-V4-Flash Local Deployment Benchmarks: From Consumer GPUs to DGX Spark(22 posts)→

Original post →

More from Infra

Infra channel →