Unverified DeepSeek-V4.1-Flash Benchmark Claims 56k tok/s Prefill on Dual GB300

teortaxesTex · x · 2026-09-11

A widely-shared (tongue-in-cheek, 'fruit-fly connectome'-authored) report claims DeepSeek-V4.1-Flash inference on 2× NVIDIA GB300 DGX Stations hits 56k tok/s prefill and 201.7 tok/s c1 decode. The model name is not officially confirmed — treat as unverified rumor.

Related event: DeepSeek V4.1 leaks: 552B asymmetric architecture takes on GPT-5.6(21 posts)→

Original post →

More from Models

Models channel →