CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 12 upvotes

REALM: A Coarse-to-Fine Generative Framework for Embodied Reactive Listening

QUESTION — How can responsive listener facial motion be generated in real-time for embodied conversational AI systems?

Generating responsive listener facial motion in conversational AI requires balancing the timing of speaker cues while maintaining motion continuity and capturing locally variable facial events. To address these challenges, the authors propose REALM (Reactive Embodied Audio-driven Listening Model), a coarse-to-fine framework combining history-aware temporal alignment with stochastic expression refinement. A Reactive Gated Speaker-Listener Fusion module combines listener motion history with speaker audio using a delay-centered attention prior and adaptive gating. A coarse decoder predicts base trajectories, augmented by audio-conditioned stochastic residuals. Evaluations on ViCo and L2L show consistent improvements over evaluated baselines across multiple motion-quality metrics.

Evaluations on ViCo and L2L show improvements over the evaluated baselines across multiple motion-quality metrics.

Deployment on an Ameca humanoid robot and a perceptual user study demonstrate the applicability of the generated behavior to physical embodiment.

Peizhen · 27 Sept 2026 read the original ↗
↑