HeteroFold Enables Prefill-Free Cross-Family KV Cache Transfer, 10.7x Faster at 32K Context
UniversityofSouthernCalifornia · hf · 2026-10-02
USC researchers propose HeteroFold, a prefill-free KV cache transfer method for heterogeneous multi-agent LLM systems, eliminating the redundant re-prefilling of shared context by receivers.
- With both sender and receiver frozen, HeteroFold handles differences in tokenization, model depth, and KV representations by aligning model structures, mapping the sender's cache into the receiver's space, and calibrating it to preserve receiver behavior
- Across six transfer directions, it achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings, matching text-based communication on a multi-agent benchmark
- At 32K context, Llama-3.1-8B → Ministral-3-14B transfer is 10.7x faster than native prefill and 1.18-1.47x faster than prior prefill-free baselines (Dense Latent, KV Ridge)
More from Infra
- Samsung reportedly quoting mid-to-high $4/Gb for HBM4, over 3x the $1.50/Gb price of HBM3E — zephyr_z9 · 2026-10-02
- Dev open-sources Bobcat, a local inference engine claiming fastest LLM runs on Apple Silicon Macs — Available_Pressure47 · 2026-10-02
- Toshiba to invest ¥60B to double AI data center HDD capacity by fiscal 2027 — zephyr_z9 · 2026-10-02
- Zuckerberg: Multi-Gigawatt Training Clusters Can Brute-Force Their Way to AGI — rohanpaul_ai · 2026-10-02
- Detect LLM hallucinations in 1.3µs on CPU — but 120B models hallucinate with unanimous false certainty — More_Slide5739 · 2026-10-02
- CodexBar: open-source menu bar app showing AI coding quotas (22k stars) — lxfater · 2026-10-02