VideoGen-Agent: Reinforcing Video Generation Agents
This paper presents VideoGen-Agent, a multimodal agent trained through multitask agentic reinforcement learning to coordinate augmentation, generation, and verification tools via multi-turn interactions. The agent undergoes supervised fine-tuning on teacher-generated trajectories to establish tool-use behaviors, followed by reinforcement learning refinement guided by a category-aware hybrid reward. Additionally, the authors introduce VABench, a held-out benchmark of 600 prompts covering identity preservation, physical consistency, and temporal structure to evaluate generation quality and tool orchestration.
On VABench, VideoGen-Agent improves over its base text-to-video generator by 19.1 points, from 56.5 to 75.6.
Upgrading the generation tools further raises the score to 86.1 without additional agent training.
Human raters prefer the upgraded configuration over the strongest standalone baseline in 84.3% of comparisons.