OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software
QUESTION — How can we systematically evaluate computer-using VLM agents on specialized scientific workflows?
The authors introduce OSWorld-Science, a benchmark and evaluation environment combining meaningful scientific tasks, artifact-based evaluation, and an efficient agent harness for computer use. The benchmark contains 12 VLMs and 146 high-quality tasks across scientific domains such as molecular drawing, pathology image analysis, statistical computing, and physical simulation. Task-specific execution-based evaluators inspect application states and generated artifacts. An integrated harness supports model adapters, interaction-loop control, and trajectory logging to facilitate model and strategy comparisons.
The benchmark contains 12 VLMs and 146 high-quality tasks across several scientific domains and software configurations.
iLOVE2D · 30 Sept 2026
read the original ↗