InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
This paper presents InterEvolve, a test-time evolution framework for humanoid loco-manipulation tasks. The system combines an object-aware forward-backward behavioral foundation model with an LLM agent that revises reward program structures in context based on execution feedback, while a numerical optimizer tunes constants. Evaluated across parallel simulation scenarios, the approach allows the system to iteratively discover and compose motor competencies, successfully executing skills autonomously on a physical Unitree G1 robot.
InterEvolve leverages an object-aware forward-backward behavioral foundation model to map rewards into loco-manipulation behaviors at test time.
An LLM agent revises the reward program structure in context using execution feedback and a skill library.
Evolved skills run autonomously on a physical Unitree G1 from egocentric onboard perception.