CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — evaluation 28 upvotes

OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

QUESTION — How can we systematically evaluate computer-using VLM agents on specialized scientific workflows?

The authors introduce OSWorld-Science, a benchmark and evaluation environment combining meaningful scientific tasks, artifact-based evaluation, and an efficient agent harness for computer use. The benchmark contains 12 VLMs and 146 high-quality tasks across scientific domains such as molecular drawing, pathology image analysis, statistical computing, and physical simulation. Task-specific execution-based evaluators inspect application states and generated artifacts. An integrated harness supports model adapters, interaction-loop control, and trajectory logging to facilitate model and strategy comparisons.

The benchmark contains 12 VLMs and 146 high-quality tasks across several scientific domains and software configurations.

iLOVE2D · 30 Sept 2026 read the original ↗
↑