DeepSeek-V4-Flash Quantized on A100: Uses Only 15.8GB VRAM at 16 tok/s
Different-Pickle1021 · reddit · 2026-08-01
A developer tested the unsloth GGUF format of DeepSeek-V4-Flash-0731 on an A100 (40GB VRAM). Using the Q8KXL quantization (162GB) with all experts offloaded to the CPU, the model achieved a generation speed of 16.1 tok/s while utilizing only 15.8GB of VRAM.
More from Infra
- Atomic-Chat: An Open-Source Local AI Assistant That Runs 100% Offline — rohanpaul_ai · 2026-08-01
- 1-bit Kimi K3 Quant Tested: 2.8T Model Compressed to 590GB Runs Locally — rohanpaul_ai · 2026-08-01
- Paper Share: How Chunked Prefill Improves LLM Serving Efficiency — Abhishekcur · 2026-08-01
- The Compute Bottlenecks of Agentic AI: Inference vs. Execution — charles_irl · 2026-08-01
- AI Market Correction Warning: Extreme Leverage in Memory Chips and Record Investor Debt — binarybits · 2026-08-01
- Micron and Hynix Cautious on Capacity Due to Memory Cycle Scars, Price Hikes Signal Expansion — _sholtodouglas · 2026-08-01