LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows
This paper presents LynnReal-Omni, a native multimodal video generation framework built on a 32B shared multimodal diffusion transformer that unifies text-to-video, image conditioning, structural control, and long-video generation from heterogeneous visual inputs. It also features a dedicated 27B Flash transformer for real-time rendering and an MSAVP evaluation design with 100 prompts and 20 metrics. On one H100, warm generation and decoding of a 22-frame 540p video take 843 ms with LynnReal-Omni and 377 ms with Flash.
LynnReal-Omni relies on a 32B shared multimodal diffusion transformer and a 27B Flash variant for real-time rendering.
Introduces MSAVP, a 100-prompt, 20-metric evaluation design separating instruction following and visual quality.
On one H100, warm generation and decoding of a 22-frame 540p video take 843 ms with LynnReal-Omni and 377 ms with Flash.