Recursive Harness Distillation across Agents for Robot Manipulation
Nghiên cứu đề xuất phương pháp Recursive Harness Distillation, cho phép một agent mạnh chắt lọc kinh nghiệm xử lý lỗi thành một cẩm nang (playbook) cho agent nhẹ, sau đó đệ quy tinh chỉnh cẩm nang này dựa trên phản hồi thực thi. Cẩm nang này giúp các agent tái sử dụng kiến thức can thiệp trong các nhiệm vụ mới mà không cần huấn luyện lại trọng số mô hình. Kết quả thực nghiệm trong môi trường thực tế và SimplerEnv Bridge cho thấy phương pháp này cải thiện đáng kể tỷ lệ thành công so với các mô hình cơ sở chỉ dùng GR00T.
In real-world manipulation, the harness improves success from 37.3% to 64.0%.
On SimplerEnv Bridge, the light agent with the playbook achieves 66.7% success, compared with 41.7% for the GR00T-only baseline, and outperforms the strong agent without a playbook.
The same playbook also benefits the strong agent, which reaches 79.2% success.