TensorFold 1.0.5 cuts 19.8k-token chat prefill from 8.1s to 0.16s with persistent prompt cache

HankYeomans · x · 2026-10-11

TensorFold 1.0.5 is out with a cross-request prompt cache for local LLM inference: later conversation turns only prefill new tokens.

Install via brew upgrade tensorfold; Living Weights coming in the next release.

Original post →

More from Infra

Infra channel →