CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — agents 49 upvotes

AutoGUIWorld: Image Generators as Visual World Models for GUI Agent

QUESTION — Can image generators serve as visual world models to synthesize interaction trajectories for GUI agents?

The authors introduce AutoGUIWorld, a data generation framework combining visual priors from image generators with a planner's task knowledge to synthesize GUI interaction trajectories without deploying actual software environments. The framework samples initial GUI scenes from OS specifications, specifies atomic actions, and iteratively edits screenshots using an image generator to produce subsequent observations. Action grounding and quality filtering yield 79,266 spatially annotated step-level training samples across Ubuntu, Windows, macOS, and Chrome. Fine-tuning Qwen3.5-35B-A3B on these trajectories improves the mean task score on OSWorld from 33.0% to 40.8% and task success rate on ScienceBoard from 14.0% to 32.2%.

AutoGUIWorld yields 79,266 spatially annotated step-level training samples across Ubuntu, Windows, macOS, and Chrome.

Fine-tuning Qwen3.5-35B-A3B on AutoGUIWorld trajectories improves the mean task score on OSWorld from 33.0% to 40.8%.

Fine-tuning Qwen3.5-35B-A3B on AutoGUIWorld trajectories improves the task success rate on ScienceBoard from 14.0% to 32.2%.

YangC777 · 01 Oct 2026 read the original ↗
↑