CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — evaluation 65 upvotes

RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

QUESTION — How can we evaluate hybrid computer-use agents that interleave graphical interactions and software implementation?

The paper investigates hybrid computer-use agents (CUAs) that autonomously interleave interface exploration, software implementation, and visual verification. The authors present RecreationWorld, a five-platform framework (Ubuntu, macOS, Windows, Android, Web) providing reproducible recreation environments with reference-grounded execution rewards. They also introduce RecreationBench, comprising 250 diverse evaluation tasks with multi-depth programmatic and visual assertions. Evaluation shows GPT-6 Astra leads at 58.1% overall, while passing all programmatic tests on only 2.8% of tasks, highlighting current limitations in interactive fidelity and computed outputs.

RecreationWorld provides reproducible environments on Ubuntu, macOS, Windows, Android, and Web with execution-grounded rewards.

RecreationBench comprises 250 diverse tasks with multi-depth programmatic and visual assertions validated by human reviewers.

GPT-6 Astra leads at 58.1% overall but passes all programmatic tests on just 2.8% of tasks.

taesiri · 18 Sept 2026 read the original ↗
↑