Unverified DeepSeek V4.1 Flash build claims 2000 tok/s prefill, 262k context
0xSero · x · 2026-10-03
User 0xSero announced a DeepSeek-V4.1-Flash-2-Sparks build (unverified by DeepSeek) with benchmark-style numbers: 2000 tok/s prefill, 29 tok/s prose decoding, 41 tok/s code decoding. The setup includes a 262k context window, a 2M-token KV cache pool, 8-way concurrency, and vision support.
Reported quality metrics: 92% top-token agreement and 0.07 KLD against the reference model, suggesting near-identical outputs. The build can be loaded via the pi tool. Authenticity remains unconfirmed.
More from Infra
- Redditor seeks RX6700XT results for Strata Qwen 3.8 Flash Next — Loose_Doubt367 · 2026-10-04
- Cerebras CEO: architecture choices sidestep HBM, CoWoS and TSMC 3nm bottlenecks — rohanpaul_ai · 2026-10-04
- Hyperscaler off-balance sheet commitments hit $3.6T, up $500B in a month — GaryMarcus · 2026-10-04
- 'One DIMM, Michael. How much could it cost, $300,000?' RAM price meme — sloppenheimer · 2026-10-04
- Report: China's growing DUVi stockpile could erase US AI chip edge within a decade; ban urged — teortaxesTex · 2026-10-04
- Compute won't follow cheap power's paradox: why Jevon's law will hold for AI compute — AccBalanced · 2026-10-04