From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention
The paper introduces PARTS (Policy Adaptation with RL on Targeted Subtasks), a real-world RL framework concentrating practice on bottlenecks of long-horizon tasks without requiring full-task demonstrations. By leveraging agent-generated selectors and success verifiers to activate residual corrections combined with online RL and success-reweighted retraining, it minimizes human intervention. On bimanual YAM and single-arm Franka robot tasks, PARTS substantially improves complete-task success rates.
On bimanual YAM and single-arm Franka tasks, PARTS improves complete-task success from 32% to 61% and from 50% to 95%, respectively, using tens of minutes of real-world RL rollouts per task on average.
Compared with existing real-world RL fine-tuning methods, PARTS raises full-task success by more than 25% under the same robot-rollout budget while requiring less human involvement.