Systematically Exploring the Capabilities of GPT-6 Astra as Embodied Policies
The study comprehensively evaluates GPT-6 Astra as a general-purpose embodied policy across six domains, examining direct control, cooperation with learned policies, and feedback-driven adaptation. Findings reveal a gap between task decisions and reliable physical control. Hybrid control with π0.5 achieves 48% success on RoboDojo and 38.7% on RoboCasa365, while navigation reaches 92% on RxR and 82% on HM3D. Inference latency heavily constrains practical control: policy-assisted and direct control consume 624.8 million and 1.132 billion tokens across 50 RoboDojo instances. A 30-second locomotion run requires 250 model calls averaging 39.86 seconds each with physics paused.
Hybrid control with π0.5 achieves 48% success on the evaluated RoboDojo subset.
In navigation, Astra reaches 92% success on RxR instruction following and 82% on HM3D object search.
A 30-second locomotion run requires 250 model calls averaging 39.86 seconds each.