CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — inference 76 upvotes

Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction

QUESTION — How can the memory footprint and cache-read traffic of the KV cache be reduced during Transformer decoding without sacrificing benchmark accuracy?

The authors introduce Grouped Value Attention (GVA), a KV cache optimization technique that stores grouped values and reconstructs content keys using a learned linear map. At inference, this map can be absorbed into the query, eliminating the need to materialize content keys in the decode path while a small shared RoPE channel retains positional information. At the 350M-parameter scale with 30B FineWeb-Edu tokens, the 16-dimensional positional variant achieves 44.18 average accuracy across five tasks, compared to 44.36 for GQA and 43.88 for MLA, while reducing persistent cache scalars by approximately 45-47% relative to matched GQA.

Grouped Value Attention (GVA) stores grouped values and reconstructs content keys using a learned linear map.

The 16-dimensional positional variant at the 350M-parameter scale with 30B FineWeb-Edu tokens achieves 44.18 average accuracy across five tasks.

GVA reduces persistent cache scalars by approximately 45-47% relative to matched GQA.

vishesh-t27 · 08 Sept 2026 read the original ↗
↑