CONSONANCE.for your information
Thứ Hai, 5 tháng 10, 2026frenvi

Đáng đọc kỹ

01 — inference 3 upvote

ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation

CÂU HỎI — Có thể loại bỏ hoàn toàn teacher-critic stack tốn kém trong việc huấn luyện hậu kỳ video tạo sinh bằng cách nào?

Nhóm tác giả nghiên cứu phương pháp ViRDM để giải quyết bài toán few-step causal video generation mà không cần dùng đến teacher và critic. Bằng cách kết hợp representation distribution matching (RDM) với stochastically truncated clean-exit supervision, lightweight VAE decoder và staged vector-Jacobian products, phương pháp này vượt qua các rào cản về bộ nhớ và thời gian. ViRDM chỉ cần 20 generator updates để đạt 84.87 trên VBench evaluation, tiêu tốn 16 A100 GPU-hours và vượt qua các baseline trước đó.

With only 20 generator updates, the recipe reaches 84.87 on the official VBench evaluation, outperforming the previous best few-step causal baseline by 0.36, while requiring 16 A100 GPU-hours.

cr8br0ze · 24 thg 9, 2026 đọc bản gốc ↗
↑