Running DeepSeek V4 Locally on 3x MI50: Hits 15 t/s

Kamal965 · reddit · 2026-08-02

A developer successfully ran the 90.9 GB DeepSeek V4-Flash-0731 model entirely in VRAM using 3x AMD MI50 GPUs (96 GB total) with the UD-IQ2M quantization format.

Performance:

Quality Test:

The author tested the model with a complex prompt to generate a 3D Rubik's Cube Canvas animation. While the model made a minor factual error (confusing memory bandwidth with PCIe 4.0 bandwidth), the author was highly impressed by the overall capability and the sheer feasibility of running such a massive model locally.

Related event: DeepSeek-V4-Flash Local Deployment Benchmarks: From Consumer GPUs to DGX Spark(22 posts)→

Original post →

More from Infra

Infra channel →