ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation
QUESTION — Can the resource-intensive teacher-critic stack in video generation post-training be completely eliminated?
The authors investigate ViRDM to address few-step causal video generation without relying on a teacher or critic network. By coupling representation distribution matching (RDM) with stochastically truncated clean-exit supervision, a lightweight VAE decoder, and staged vector-Jacobian products, the method overcomes memory and gradient bottlenecks. ViRDM requires only 20 generator updates to reach 84.87 on the official VBench evaluation, consuming 16 A100 GPU-hours while outperforming previous causal baselines.
With only 20 generator updates, the recipe reaches 84.87 on the official VBench evaluation, outperforming the previous best few-step causal baseline by 0.36, while requiring 16 A100 GPU-hours.
cr8br0ze · 24 Sept 2026
read the original ↗