Unverified DeepSeek-V4.1-Flash Benchmark Claims 56k tok/s Prefill on Dual GB300
teortaxesTex · x · 2026-09-11
A widely-shared (tongue-in-cheek, 'fruit-fly connectome'-authored) report claims DeepSeek-V4.1-Flash inference on 2× NVIDIA GB300 DGX Stations hits 56k tok/s prefill and 201.7 tok/s c1 decode. The model name is not officially confirmed — treat as unverified rumor.
Related event: DeepSeek V4.1 leaks: 552B asymmetric architecture takes on GPT-5.6(21 posts)→
More from Models
- Falcon OCR 1.5: a 270M model challenges OCR rivals 5-20x its size — HildeKuehne · 2026-09-11
- Switching from Claude to Codex Pro: faster but sloppier, loses context — green_r00t · 2026-09-11
- Shisa.AI publishes nuclear-design and multilingual multi-agent use cases for its models — bclavie · 2026-09-11
- User reports new Opus coding safety filters refuse biology requests outright — iskander · 2026-09-11
- DeepSeek V4.1 Flash reportedly bakes prefill/decode disaggregation into the model weights — altryne · 2026-09-11
- Yemeni terrorists reportedly used Claude Code to program rockets, sparking misuse debate — davidmanheim · 2026-09-11