CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — inference 67 upvotes

Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs

QUESTION — How can KV cache be transferred across heterogeneous model families in multi-agent systems without requiring receiver prefill?

The authors propose HeteroFold, a prefill-free cross-family KV cache transfer method that keeps both sender and receiver models frozen. By aligning model structures, mapping the sender cache into the receiver space, and calibrating it, HeteroFold enables efficient cross-family KV reuse without receiver prefill. Across evaluations, the method achieves top cache-transfer performance on all four long-context benchmarks and matches text-based communication. At a 32K context length, the Llama-3.1-8B to Ministral-3-14B transfer achieves a speedup of 10.7 times compared to Native Prefill and 1.18 to 1.47 times faster than state-of-the-art baselines like Dense Latent and KV Ridge.

HeteroFold achieves the best cache-transfer performance on all four long-context benchmarks and most short-context settings.

At 32K context length, Llama-3.1-8BrightarrowMinistral-3-14B transfer is 10.7times faster than Native Prefill.

It is 1.18--1.47times faster than the state-of-the-art prefill-free baselines, Dense Latent and KV Ridge.

yunuyean · 29 Sept 2026 read the original ↗
↑