All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation
QUESTION — How can the cross-modal correspondence asymmetry between video and companion modalities be resolved in joint multimodal diffusion transformers?
The study reveals an asymmetry in joint multimodal diffusion transformers: companion modalities like 3D body motion or audio form strong correspondences to video, but the reverse correspondences are weaker, a discrepancy termed the reciprocal correspondence gap. The authors introduce RecCAR (Reciprocal Cross-modal Attention Regularization), a KL regularizer that uses established video-to-modality correspondence as a fixed reference to align the weaker modality-to-video correspondence. Experiments across video-motion and video-audio generation show that RecCAR improves anatomical scores and reduces audiovisual desynchronization.
Improves the Human Anatomy score from 0.69 to 0.75.
Reduces audio-video desynchronization from 0.804 to 0.752.
ohad204 · 23 Sept 2026
read the original ↗