DeepSeek-V4 Local Test: 17 t/s on A6000 + 256GB RAM

USBhost · reddit · 2026-08-02

A developer on Reddit shared local inference benchmarks for DeepSeek-V4-Flash-0731 (Q8 quantization). Running on an RTX A6000 (48GB) paired with 256GB of DDR4 RAM, the model achieved a steady generation speed of 17.2 tokens/s and prompt processing over 70 tokens/s. The author noted that while 48GB VRAM is theoretically sufficient for a 1 million token context, prompt processing speeds would drop significantly.

Related event: DeepSeek-V4-Flash Local Deployment Benchmarks: Performance Across Hardware(21 posts)→

Original post →

More from Infra

Infra channel →