Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms
This work investigates the root cause of physical law violations in video diffusion models by analyzing their internal attention mechanisms. The authors present an interpretability study on the motion planning process, revealing that Rotary Position Embedding (RoPE) induces excessive spatial attention decay, causing early candidate regions to lock into physically implausible positions. To address this flaw, they propose a lightweight architectural modification that scales the frequency of RoPE across different denoising steps to reduce attention decay and improve physical commonsense.
Rotary Position Embedding (RoPE) induces excessive spatial attention decay that triggers generation failure modes.
A lightweight architectural modification scaling the frequency of RoPE across denoising steps enhances physical commonsense.