Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation
This paper addresses the gap in evaluating multi-step physical reasoning within egocentric video generation models for embodied planning. The authors introduce Ego2Act, a benchmark featuring 2,640 videos spanning 110 real-world tasks with varying clutter and multi-step complexity, alongside Ego2ActJudge, a reference-free evaluation pipeline aligned with human consensus for task completion and physics plausibility. Experimental findings reveal that current models frequently skip or partially execute steps and consistently fail at fine-grained physical dynamics during complex object manipulation.
Ego2Act features 2,640 videos from 110 real-world tasks across day-to-day settings, varying object clutter and multi-step complexity.
Ego2ActJudge achieves better task completion and physics plausibility evaluation alignment with human consensus compared to relevant baselines.