CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 8 upvotes

All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation

QUESTION — How can the cross-modal correspondence asymmetry between video and companion modalities be resolved in joint multimodal diffusion transformers?

The study reveals an asymmetry in joint multimodal diffusion transformers: companion modalities like 3D body motion or audio form strong correspondences to video, but the reverse correspondences are weaker, a discrepancy termed the reciprocal correspondence gap. The authors introduce RecCAR (Reciprocal Cross-modal Attention Regularization), a KL regularizer that uses established video-to-modality correspondence as a fixed reference to align the weaker modality-to-video correspondence. Experiments across video-motion and video-audio generation show that RecCAR improves anatomical scores and reduces audiovisual desynchronization.

Improves the Human Anatomy score from 0.69 to 0.75.

Reduces audio-video desynchronization from 0.804 to 0.752.

ohad204 · 23 Sept 2026 read the original ↗
↑