CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — multimodal 57 upvotes

EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos

QUESTION — How well do current multimodal video models perform on tool-centric reasoning in real-world egocentric videos?

Real-world tasks require agents to act under physical constraints where tool use is central. However, current multimodal video models remain limited in tool-centric reasoning due to a lack of real-world egocentric data and diagnostic benchmarks. This study introduces EgoTools, consisting of EgoTools-Data (a 100-hour egocentric recording corpus) and EgoTools-Bench (a diagnostic benchmark of 1,000 QA pairs). Experimental results show that current models still struggle to ground tool use in visual evidence, but full supervised fine-tuning with EgoTools-Data significantly improves model performance.

Gemini-3.1-Pro achieves 66.9% overall accuracy but only 51.7% on Perception & Grounding.

Full supervised fine-tuning improves Qwen3-VL-8B-Instruct from 50.0% to 60.9% on the full 1,000-question benchmark.

shulin16 · 30 Sept 2026 read the original ↗
↑