OmniEcho: Spatial Audio Understanding for Embodied Agents
This research addresses spatial audio evaluation gaps for embodied agents by introducing OmniEchoBench, a benchmark featuring six tasks across 197 real-world audio-visual scenes, 2,972 Q&A pairs, and 900 navigation samples collected from 30 environments using first-order ambisonics (FOA) audio. The authors also develop a controllable rendering pipeline and propose OmniEcho, a spatially aware omni-modal model featuring an FOA spatial encoder. Experiments show state-of-the-art perception performance and navigation capabilities close to traditional vision-language models.
OmniEchoBench features six tasks across 197 real-world scenes with 2,972 Q&A pairs and 900 navigation samples.
Data uses first-order ambisonics (FOA) audio collected from 30 real-world environments.
OmniEcho incorporates an FOA spatial encoder to achieve state-of-the-art spatial audio-visual perception.