Peking Univ. Releases TensorCast: 228x Faster Cold Starts, 93.2% Lower TTFT
jiqizhixin · x · 2026-08-24
Peking University, StepFun, and BUPT introduce TensorCast, a unified programmable tensor management layer between compute, network, and storage. It decouples tensor states (weights, KV cache, checkpoints) from specific systems. Results show it matches Mooncake's KV cache performance, cuts instance startup time by up to 228.6x for 235B-parameter models, and reduces TTFT by 93.2% in high-concurrency multi-turn agent scenarios.
More from Infra
- Stanford's Marin 535B Model Training Starts with Full Transparency — udmrzn · 2026-08-24
- Seeking Real-World Benchmarks: R9 7900XTX Running Qwen2.5-72B — BillyQ · 2026-08-24
- Splitting GPUs Across VMs Boosted Ollama Performance by ~3x — NicolaZanarini533 · 2026-08-24
- Gated DeltaNet-2 gets full cuDNN support, ~3x faster end-to-end training on NVIDIA GPUs — ZGojcic · 2026-08-24
- M5 Max vs RTX 5080/5090 for local visual AI workloads — durumertt · 2026-08-24
- Hippius launches decentralized storage at 1/100th of Big Cloud costs — markjeffrey · 2026-08-24