Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction
The authors introduce Grouped Value Attention (GVA), a KV cache optimization technique that stores grouped values and reconstructs content keys using a learned linear map. At inference, this map can be absorbed into the query, eliminating the need to materialize content keys in the decode path while a small shared RoPE channel retains positional information. At the 350M-parameter scale with 30B FineWeb-Edu tokens, the 16-dimensional positional variant achieves 44.18 average accuracy across five tasks, compared to 44.36 for GQA and 43.88 for MLA, while reducing persistent cache scalars by approximately 45-47% relative to matched GQA.
Grouped Value Attention (GVA) stores grouped values and reconstructs content keys using a learned linear map.
The 16-dimensional positional variant at the 350M-parameter scale with 30B FineWeb-Edu tokens achieves 44.18 average accuracy across five tasks.
GVA reduces persistent cache scalars by approximately 45-47% relative to matched GQA.