WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
QUESTION — How can long-horizon consistency and camera-control accuracy be maintained in video world models?
The authors introduce WorldCrafter, a video world model integrating a camera-queryable implicit 3D-aware memory to handle long-horizon observations. The system employs a memory encoder and a pose-conditioned readout module to aggregate historical observations into fixed target tokens prior to denoising. By combining this memory with temporal context and few-step distillation, the model enables streaming scene exploration from a single input image or text prompt.
WorldCrafter learns a camera-queryable implicit 3D-aware memory to improve long-horizon consistency.
Drexubery · 21 Sept 2026
read the original ↗