CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 25 upvotes

Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

QUESTION — How can goal-directed manipulation in egocentric video generation models be rigorously evaluated?

This paper addresses the gap in evaluating multi-step physical reasoning within egocentric video generation models for embodied planning. The authors introduce Ego2Act, a benchmark featuring 2,640 videos spanning 110 real-world tasks with varying clutter and multi-step complexity, alongside Ego2ActJudge, a reference-free evaluation pipeline aligned with human consensus for task completion and physics plausibility. Experimental findings reveal that current models frequently skip or partially execute steps and consistently fail at fine-grained physical dynamics during complex object manipulation.

Ego2Act features 2,640 videos from 110 real-world tasks across day-to-day settings, varying object clutter and multi-step complexity.

Ego2ActJudge achieves better task completion and physics plausibility evaluation alignment with human consensus compared to relevant baselines.

patrickamadeus · 01 Oct 2026 read the original ↗
↑