7x RTX 3090s Barely Run DeepSeek-V4 Q8 Quantization Locally

_akhaliq · x · 2026-08-01

A developer shared that their rig of 7x RTX 3090s can just barely run the Q8 quantized version of unsloth/DeepSeek-V4-Flash-0731-GGUF. This demonstrates that with community quantization, consumer-grade hardware clusters are now on the edge of running massive inference models locally.

Original post →

More from Infra

Infra channel →