CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — agents 55 upvotes

InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation

QUESTION — How can reward programs evolve at test time to repurpose humanoid control skills without retraining?

This paper presents InterEvolve, a test-time evolution framework for humanoid loco-manipulation tasks. The system combines an object-aware forward-backward behavioral foundation model with an LLM agent that revises reward program structures in context based on execution feedback, while a numerical optimizer tunes constants. Evaluated across parallel simulation scenarios, the approach allows the system to iteratively discover and compose motor competencies, successfully executing skills autonomously on a physical Unitree G1 robot.

InterEvolve leverages an object-aware forward-backward behavioral foundation model to map rewards into loco-manipulation behaviors at test time.

An LLM agent revises the reward program structure in context using execution feedback and a skill library.

Evolved skills run autonomously on a physical Unitree G1 from egocentric onboard perception.

xusirui · 01 Oct 2026 read the original ↗
↑