TensorFold 0.6.1: 36% faster 27B inference, first token drops to 0.1s on Blackwell

EAccelerate_42 · x · 2026-10-02

TensorFold 0.6.1 ships major inference optimizations: 27B models run directly on NVIDIA RTX PRO 6000 Blackwell with NVFP4 4-bit checkpoints; Flash Next on CUDA enables one-pass prompt filling, up to 36% faster throughput with 8 streams, and slowest first-token latency cut from 0.8s to 0.1s. A shared 40,900-token system prompt now answers in under a second instead of 16-56s; forks resume from split points, short requests no longer queue, and images work alongside text. DGX Spark prompts fill 15% faster; Macs get up to 10%. The author notes speculative decoding and TensorFold dramatically boost DGX Spark's value.

Original post →

More from Infra

Infra channel →