AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation
QUESTION — How can we resolve entangling learning signals in reinforcement learning optimization for joint audio-video generation?
Joint audio-video generation struggles because heterogeneous multimodal rewards entangle learning signals, and optimizing two modality towers is computationally expensive. To address this, the authors propose AV-GRPO, a modality-anchored online diffusion RL framework, along with the 5DAV dataset. AV-GRPO features modality-anchored rollouts, trajectory-locked frozen-tower optimization, and adaptive objectives to convert coupled multimodal preference learning into unimodal subproblems. Experiments on JavisBench and VABench demonstrate superior generation quality and synchronization compared to LTX-2.3.
AV-GRPO outperforms LTX-2.3 in generation quality, semantic alignment and cross-modal synchronization under LoRA and full fine-tuning.
Dr-Loser · 24 Sept 2026
read the original ↗