WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms
KV cache memory and bandwidth costs limit efficient long-context inference. The authors introduce WUSH-KV for low-bit KV-cache quantization, adapting WUSH to construct a data-aware transform from second-order statistics of matrix products. WUSH-KV uses calibration data to construct separate key and value transforms, folding the value transform into model weights and applying the key transform after RoPE. Integrated into SGLang using OSCAR-style percentile-clipped affine quantization, WUSH-KV reduces layerwise reconstruction error and achieves the lowest end-to-end perplexity. At 2-bit, WUSH-KV performs comparably to or outperforms the OSCAR transform across evaluated models and tasks.
WUSH-KV uses calibration data to construct separate key and value transforms.
The value transform is folded into the model weights and the key transform is applied after RoPE.
At 2-bit, WUSH-KV performs comparably to or outperforms the OSCAR transform across all evaluated models and tasks.