DeepSeek-V4-Flash Quantized on A100: Uses Only 15.8GB VRAM at 16 tok/s

Different-Pickle1021 · reddit · 2026-08-01

A developer tested the unsloth GGUF format of DeepSeek-V4-Flash-0731 on an A100 (40GB VRAM). Using the Q8KXL quantization (162GB) with all experts offloaded to the CPU, the model achieved a generation speed of 16.1 tok/s while utilizing only 15.8GB of VRAM.

Original post →

More from Infra

Infra channel →