Memorizon: Training World Models Beyond Their Context Window
The paper introduces Memorizon, a training method that decouples long-span supervision from expensive dense attention for streaming world models. Instead of computing attention over the full history, scored chunks retrieve top-K latents by camera co-visibility into a bounded shared bank (kK), keeping sequence costs stable regardless of span length. Moving from 100 to 400 seconds adds only 12% to the step time. Experiments show that retrieval significantly raises revisit consistency across all evaluation splits, and filling the bank from another episode lowers revisit correlation by 83%.
Going from 100 to 400 s adds 12% to the step time.
Against a sliding-window baseline, retrieval raises revisit consistency on every split.
Filling the bank from another episode lowers revisit correlation by 83%, so the model uses what it retrieves.