TensorFold 0.6.1: 36% faster 27B inference, first token drops to 0.1s on Blackwell
EAccelerate_42 · x · 2026-10-02
TensorFold 0.6.1 ships major inference optimizations: 27B models run directly on NVIDIA RTX PRO 6000 Blackwell with NVFP4 4-bit checkpoints; Flash Next on CUDA enables one-pass prompt filling, up to 36% faster throughput with 8 streams, and slowest first-token latency cut from 0.8s to 0.1s. A shared 40,900-token system prompt now answers in under a second instead of 16-56s; forks resume from split points, short requests no longer queue, and images work alongside text. DGX Spark prompts fill 15% faster; Macs get up to 10%. The author notes speculative decoding and TensorFold dramatically boost DGX Spark's value.
More from Infra
- Gemini 4 may roll out to Ultra users next week as Google's TPU capacity reportedly runs tight — haider1 · 2026-10-02
- Banks tighten data center lending as Beff bets thermodynamic computing can fix AI power crunch — beffjezos · 2026-10-02
- Recovering 4.1 GiB of hidden RAM on NVIDIA DGX Spark, 3.6x bigger KV pool — marian_nmt · 2026-10-02
- Local Qwen 27B 4-bit proved its worth in a hospital with no signal or Wi-Fi — AaronBergman18 · 2026-10-02
- Open-weight models trail frontier by just 4 months — here's when to use them — TechPreacher · 2026-10-02
- Google's first space compute prototype satellite launches on SpaceX rocket — elonmusk · 2026-10-02