539 tok/s DeepSeek on 4x RTX 6000 — and a call-out that community benchmarks inflate 20-30%

HankYeomans · x · 2026-10-05

A local inference enthusiast recorded 539 tok/s decode (518 median) running DeepSeek V41 Flash on 4x RTX 6000 with TP4 and stable OC (+350 core / +6000 mem).

Beyond the numbers, the poster calls out the community: there's no standard measurement for decode/prefill, and many headline tok/s figures turn out inflated 20-30% on top of using highly predictable prompts plus draft-model speculative decoding that alone adds 100 tok/s. Most people can't verify with 4x 6000s, so inflated numbers go unnoticed. He urges the local AI community to establish standard benchmarks and question all published numbers, including his own.

Original post →

More from Infra

Infra channel →