TensorFold 1.0.5 cuts 19.8k-token chat prefill from 8.1s to 0.16s with persistent prompt cache
HankYeomans · x · 2026-10-11
TensorFold 1.0.5 is out with a cross-request prompt cache for local LLM inference: later conversation turns only prefill new tokens.
- Nemotron 3.5 Lightning: a 19.8k-token chat's second turn prefills in 0.16s instead of 8.1s on M3 Ultra; on M5 Max, later turns of a 7k-token chat hit first token in 0.02-0.17s instead of 1.6-2.1s
- GLM-5.3-Flash now runs on a single 256 GB M5 Ultra
- Editing/deleting/regenerating the last turn also resumes from a cached state (9k-token chat: 1.7s vs 10.9s total prompt processing on M3 Ultra)
- --prompt-cache-gib sizes the cache, 0 disables it; fixed a 27B issue on macOS 27
Install via brew upgrade tensorfold; Living Weights coming in the next release.
More from Infra
- AirLLM runs 70B models on a 4GB GPU via layer-wise inference, scaling to 405B on 8GB — JensHonack · 2026-10-11
- Speech Model Shrunk 13x to 153M Params by Looping 2 Shared Blocks — pbaylies · 2026-10-11
- Inference demand went vertical, yet is a flat line next to post-training/RL growth — zainhas · 2026-10-11
- Pat Gelsinger slams HBM as "a lousy memory" wasting four bits for every one it makes — SumitGup · 2026-10-11
- Qualcomm CEO predicts AI phone supercycle, smart glasses as top AI wearable — SuB8u · 2026-10-11
- Zero cold starts: Building and shipping MCP servers with WebAssembly, Spin and Akamai Functions — AI Engineer · 2026-10-11