539 tok/s DeepSeek on 4x RTX 6000 — and a call-out that community benchmarks inflate 20-30%
HankYeomans · x · 2026-10-05
A local inference enthusiast recorded 539 tok/s decode (518 median) running DeepSeek V41 Flash on 4x RTX 6000 with TP4 and stable OC (+350 core / +6000 mem).
Beyond the numbers, the poster calls out the community: there's no standard measurement for decode/prefill, and many headline tok/s figures turn out inflated 20-30% on top of using highly predictable prompts plus draft-model speculative decoding that alone adds 100 tok/s. Most people can't verify with 4x 6000s, so inflated numbers go unnoticed. He urges the local AI community to establish standard benchmarks and question all published numbers, including his own.
More from Infra
- Muse ships reliability fixes after SEVs left cron jobs and scheduled tasks unrecovered — alexandr_wang · 2026-10-05
- AWS Mistakenly Suspends Account, Wabi Down for 3+ Hours With No Recourse — soleio · 2026-10-05
- Is Strix Halo the closest thing to a dream local LLM box? Unified memory vs GPUs for 20B-32B models — Robert__Sinclair · 2026-10-05
- GLM 5.3 flash on dual DGX Sparks gets 50-90% decode boost with new open recipe — swiebertjee · 2026-10-05
- Qualcomm's Snapdragon to power next-gen AI assistants for Meta and OpenAI — ryanshrout · 2026-10-05
- Spite: a modular Rust inference engine where every model, GPU and op is pluggable — giveen · 2026-10-05