Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
The authors propose HeteroFold, a prefill-free cross-family KV cache transfer method that keeps both sender and receiver models frozen. By aligning model structures, mapping the sender cache into the receiver space, and calibrating it, HeteroFold enables efficient cross-family KV reuse without receiver prefill. Across evaluations, the method achieves top cache-transfer performance on all four long-context benchmarks and matches text-based communication. At a 32K context length, the Llama-3.1-8B to Ministral-3-14B transfer achieves a speedup of 10.7 times compared to Native Prefill and 1.18 to 1.47 times faster than state-of-the-art baselines like Dense Latent and KV Ridge.
HeteroFold achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings.
At 32K context length, Llama-3.1-8BrightarrowMinistral-3-14B transfer is 10.7times faster than Native Prefill.
It is 1.18--1.47times faster than the state-of-the-art prefill-free baselines, Dense Latent and KV Ridge.