WorldAttention: An Efficient Attention Architecture for Interactive Video World Models
QUESTION — What attention architecture can resolve KV cache memory saturation and computational complexity in video world models?
This research introduces WorldAttention, a system-oriented attention architecture designed for interactive video world models. By co-designing specialized attention kernels and hierarchical memory management, it features Hybrid Sparse Attention (HSA) combining linear global attention with head-adaptive sparse attention, alongside a Hierarchical KV Cache (HKV) that organizes historical KV pairs across multi-tier memory. This design successfully eliminates GPU memory saturation and computational overhead while maintaining historical context.
WorldAttention consistently surpasses prior state-of-the-art methods, achieving subject consistency scores of 0.9472 on VBench-Long and 0.9668 on InterVBench, respectively.
SteveZeyuZhang · 28 Sept 2026
read the original ↗