Running DeepSeek V4 Locally on 3x MI50: Hits 15 t/s
Kamal965 · reddit · 2026-08-02
A developer successfully ran the 90.9 GB DeepSeek V4-Flash-0731 model entirely in VRAM using 3x AMD MI50 GPUs (96 GB total) with the UD-IQ2M quantization format.
Performance:
- Text Generation: Stable at 15-16 tokens/second, never dropping below 14 t/s even during a 30K token output.
- Prompt Processing: Around 105-110 tokens/second.
Quality Test:
The author tested the model with a complex prompt to generate a 3D Rubik's Cube Canvas animation. While the model made a minor factual error (confusing memory bandwidth with PCIe 4.0 bandwidth), the author was highly impressed by the overall capability and the sheer feasibility of running such a massive model locally.
More from Infra
- AMD MI355X Beats NVIDIA B200 in Kimi K3 Deployment with 952 tok/s — adrianscottcom · 2026-08-03
- MiniMax H3 Gets Day 0 Support in SGLang, Runs Locally on Dual RTX 5090s — ying11231 · 2026-08-03
- Qwen3.8-27B Open Weights Coming, Runs Locally on 17GB RAM — danielhanchen · 2026-08-03
- AirLLM Breaks VRAM Barrier: Runs 70B LLMs on a Single 4GB GPU — techNmak · 2026-08-03
- MiniMax H3 Open Weights Hit fal with Out-of-the-Box Inference Optimizations — gorkem · 2026-08-03
- ComfyUI Adds Day 0 Support for MiniMax Video Model, Slashing VRAM by 66% for RTX 3060 — crystal_alpine · 2026-08-03