Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention
QUESTION — How can the expressivity of delta-rule-based Linear RNNs be enhanced by combining parameter range extensions for rotations and reflections?
This research investigates enhancing the expressivity of Kimi Delta Attention (KDA) by combining a single delta-rule transformation with a second reflection supplied by its channel-wise gate. The authors extend KDA's parameter ranges by allowing gates in [-1,1] and the delta-rule coefficient β in [0,2], resulting in Complex KDA (CKDA). CKDA preserves stability and efficiency with diagonal-plus-rank-one transitions while reaching the state-tracking expressivity of DeltaProduct_2. Experiments show that CKDA outperforms Transformers and other linear RNNs in language modeling and achieves strong length extrapolation.
Every orthogonal diagonal-plus-rank-one matrix is exactly a CKDA transition matrix.
korbip · 21 Sept 2026
read the original ↗