CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — inference 49 upvotes

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

QUESTION — How can a hybrid attention mechanism combining local Softmax and bidirectional linear memory accelerate livestream video generation while preserving fine-grained quality?

The authors present Video DeltaNet (VDN), a hybrid attention architecture designed to alleviate computational bottlenecks in video diffusion models. VDN's linear branch introduces Video Delta Attention to update memory per frame using spatial tokens, while keeping Softmax attention for text and audio interactions. Utilizing a staged teacher-alignment recipe, an eight-step distillation, and an optimized SGLang serving stack, VDN-H3 efficiently completes denoising for long high-resolution video streams on multiple NVIDIA B200 GPUs.

VDN-H3 completes DiT denoising for a 14.3-second, 768p video in 6.70 seconds on eight NVIDIA B200 GPUs.

The approach achieves a 14.5x speedup over the 50-step dense H3 baseline on the same GPU count.

The architecture applies an eight-step distillation combined with an optimized SGLang serving stack.

taesiri · 17 Sept 2026 read the original ↗
↑