MintAct: A Unified Visual Agent for Digital Environments
QUESTION — How can a unified vision-language model family be trained to handle UI grounding, multi-step navigation, and visual tool use across diverse digital environments?
MintAct is a family of vision-language models at 2B, 4B, and 8B scales that unifies UI grounding, multi-step navigation across mobile, desktop, and web, and visual tool use. To train these models, the authors developed a scalable environment and reinforcement learning infrastructure hosting hundreds of concurrent instances across heterogeneous backends. An asynchronous framework maintains control over the cross-domain training distribution and remains stable under noisy feedback. Experimental results show that MintAct achieves state-of-the-art performance on OSWorld-Verified with a score of 48.9.
MintAct achieves state-of-the-art performance (48.9 on OSWorld-Verified) across a wide range of benchmarks at comparable model sizes.
taesiri · 18 Sept 2026
read the original ↗