DeepSeek-V4.1-Flash on 2x DGX Spark: TensorFold 1.0 doubles 128K decode to 100 tok/s

EAccelerate_42 · x · 2026-10-09

A v0.6.0 release of the Zig-native TensorFold engine runs DeepSeek-V4.1-Flash on 2x DGX Spark with major speedups: 128K-deep decode nearly doubles from 49.8 to 100.1 tok/s, 4-stream sustained output rises 13% to 142.0 tok/s, and 128K prefill gains 5%, while supporting 1M context plus vision with bit-for-bit identical outputs. Code is on GitHub. Note: the model name is not an official DeepSeek release; treat as unverified third-party claim.

Original post →

More from Infra

Infra channel →