WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
The paper presents WanPE, a 397B-parameter prompt enhancement model trained on 1.05M real-world videos to handle shot-level cinematic planning. WanPE combines video-grounded reverse construction with Semantic-Consistency GRPO (SC-GRPO) to maintain user constraints across multiple shots and time. When powering Wan3.0, WanPE significantly boosts human preference scores over raw prompts, showing dramatic performance gains especially in 30-second video generation benchmarks based on blind pairwise assessments.
WanPE is a 397B-parameter prompt enhancement model trained on 1.05M videos.
WanPE employs Semantic-Consistency GRPO (SC-GRPO) to preserve semantic fidelity.
WanPE boosts human preference scores by 10.66-18.84 points at 5-15 seconds and 50.86 points at 30 seconds.