OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning
The paper presents OmniSeek, an agentic framework that transforms an Omni Large Language Model into an active, multi-turn reasoning agent with native tool use. Instead of passively processing an entire audio-visual sequence in a single forward pass, OmniSeek makes evidence acquisition part of the reasoning process by dynamically deciding whether to look or listen over temporal windows. They build a data engine synthesizing OmniTraj-170K containing multi-hop Chain-of-Thought trajectories, supervise the model on these, and optimize the policy via two-stage reinforcement learning with verifiable rewards. Experiments show OmniSeek improves audio-visual reasoning performance.
OmniSeek transforms an Omni-LLM into a multi-turn reasoning agent with native tool use.
A data engine synthesizes OmniTraj-170K, a corpus of multi-hop Chain-of-Thought trajectories with interleaved audio and visual evidence.
A two-stage reinforcement learning with verifiable rewards and an Audio-Visual Necessity objective optimizes the policy.