TensorCast Unifies Tensor Management, Boosting LLM Startup Speed by 228x

机器之心 · wechat · 2026-08-17

Peking University, in collaboration with StepFun, proposes TensorCast, a unified abstraction layer for tensor lifecycle management in LLM infrastructure. By decoupling states like weights and KVCache from specific compute systems, it enables programmable cross-component optimization strategies. Experiments show TensorCast boosts model startup speed by up to 228.6x, matches specialized KVCache systems in performance, and reduces median TTFT by 93.2% in high-concurrency multi-turn Agent workloads.

Original post →

More from Infra

Infra channel →