DeepSeek-V4-Flash tops out at 770 tok/s on one B300 in a vLLM batch test

Moreh · reddit · 2026-07-22

A Reddit user is benchmarking DeepSeek-V4-Flash on a single B300 with vLLM 0.25.0 and says throughput is only about 770 output tokens/s at batch size 256.

Original post →

More from Infra

Infra channel →