DeepSeek-V4.1-Flash on 2x DGX Spark: TensorFold 1.0 doubles 128K decode to 100 tok/s
EAccelerate_42 · x · 2026-10-09
A v0.6.0 release of the Zig-native TensorFold engine runs DeepSeek-V4.1-Flash on 2x DGX Spark with major speedups: 128K-deep decode nearly doubles from 49.8 to 100.1 tok/s, 4-stream sustained output rises 13% to 142.0 tok/s, and 128K prefill gains 5%, while supporting 1M context plus vision with bit-for-bit identical outputs. Code is on GitHub. Note: the model name is not an official DeepSeek release; treat as unverified third-party claim.
More from Infra
- Amazon drops data center NDAs as community backlash spurs hundreds of moratoriums — TechCrunch AI · 2026-10-10
- Amazon drops data center NDAs, and AI agents want your credit card — TechCrunch AI · 2026-10-10
- SkyPilot founder: hoarded idle GPUs waste $20M+ a year, and the AI compute layer should be open — skypilot_org · 2026-10-10
- uv binary shrinks over 40% since July, saving nearly 4 PB of mirror bandwidth monthly — charliermarsh · 2026-10-10
- Firmus' $30B IPO collapses: only 46MW operational, valuation tripled in two months — kevinsxu · 2026-10-10
- Hyperscaler bonds now a notable share of net new Treasury borrowing — matt_slotnick · 2026-10-10