RoboFoundry: System-as-Policy Evolution for Self-Learning Embodied Agents
The authors propose RoboFoundry, an embodied agentic framework formulating self-improvement as a Self-Evolving System-as-Policy. The framework diagnoses capability gaps and converts execution traces into validated system updates across an active context system and a hierarchical skill system. On EmbodiedBench, RoboFoundry improves GPT-5.5 by 27.8% and brings Qwen3.7-Plus to near parity at 70.3% versus 72.7%. Furthermore, it outperforms baselines on RoboMemArena by at least 39.0% and surpasses Cap-Agent0 by 243.8%-679.7% on LIBERO-PRO.
On EmbodiedBench, RoboFoundry achieves state-of-the-art performance, notably improving GPT-5.5 by 27.8%.
It also brings Qwen3.7-Plus to near parity with GPT-5.5 (70.3% vs. 72.7%), showing consistent gains from system-as-policy evolution across foundation models.
For long-horizon memory, RoboFoundry outperforms all baselines on RoboMemArena by at least 39.0%, even against methods assisted by external foundation models.