LOCI: Spatial Linear Memory for Streaming World Models
QUESTION — How can we maintain long-term spatial memory for streaming video world models while controlling peak memory consumption?
LOCI is a hybrid spatial-memory architecture for streaming video world models. It splits transformer blocks between those maintaining a key-value cache of past observations and those using chunk-restricted attention complemented by a recurrent linear-attention memory conditioned on projective camera geometry. Experiments on the MIND benchmark and recorded trajectories show that LOCI reconstructs revisited regions more faithfully than baseline models, lowering peak memory by about 30% relative to full softmax at equal length.
With full history, it lowers peak memory at equal length by about 30% relative to full softmax.
sum0214 · 30 Sept 2026
read the original ↗