AutoGUIWorld: Image Generators as Visual World Models for GUI Agent
The authors introduce AutoGUIWorld, a data generation framework combining visual priors from image generators with a planner's task knowledge to synthesize GUI interaction trajectories without deploying actual software environments. The framework samples initial GUI scenes from OS specifications, specifies atomic actions, and iteratively edits screenshots using an image generator to produce subsequent observations. Action grounding and quality filtering yield 79,266 spatially annotated step-level training samples across Ubuntu, Windows, macOS, and Chrome. Fine-tuning Qwen3.5-35B-A3B on these trajectories improves the mean task score on OSWorld from 33.0% to 40.8% and task success rate on ScienceBoard from 14.0% to 32.2%.
AutoGUIWorld yields 79,266 spatially annotated step-level training samples across Ubuntu, Windows, macOS, and Chrome.
Fine-tuning Qwen3.5-35B-A3B on AutoGUIWorld trajectories improves the mean task score on OSWorld from 33.0% to 40.8%.
Fine-tuning Qwen3.5-35B-A3B on AutoGUIWorld trajectories improves the task success rate on ScienceBoard from 14.0% to 32.2%.