OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
QUESTION — How can native audio-visual dialogue models be trained and evaluated when real-world data is scarce?
The research defines the OmniVChat task for direct audio-visual dialogue and introduces OmniVChat-Studio, a multi-agent data engine to synthesize single- and multi-turn dialogues. They build OmniVChat-Bench for evaluating basic capabilities and design the OmniVChat-RL reinforcement learning reward targeting reply correctness, efficiency, and style. Training Qwen3-Omni-Instruct with this framework improves performance on both synthesized benchmarks and human-recorded datasets.
OmniVChat-Studio synthesizes single- and multi-turn audio-visual dialogues for training and evaluation.
Harland · 18 Sept 2026
read the original ↗