CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — agents 15 upvotes

OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning

QUESTION — How does OmniSeek transform an Omni-LLM into a multi-turn audio-visual reasoning agent?

The paper presents OmniSeek, an agentic framework that transforms an Omni Large Language Model into an active, multi-turn reasoning agent with native tool use. Instead of passively processing an entire audio-visual sequence in a single forward pass, OmniSeek makes evidence acquisition part of the reasoning process by dynamically deciding whether to look or listen over temporal windows. They build a data engine synthesizing OmniTraj-170K containing multi-hop Chain-of-Thought trajectories, supervise the model on these, and optimize the policy via two-stage reinforcement learning with verifiable rewards. Experiments show OmniSeek improves audio-visual reasoning performance.

OmniSeek transforms an Omni-LLM into a multi-turn reasoning agent with native tool use.

A data engine synthesizes OmniTraj-170K, a corpus of multi-hop Chain-of-Thought trajectories with interleaved audio and visual evidence.

A two-stage reinforcement learning with verifiable rewards and an Audio-Visual Necessity objective optimizes the policy.

WHB139426 · 01 Oct 2026 read the original ↗
↑