CompoWorld: Compositional Environment Scaling for General Agents
The authors introduce CompoWorld, a framework that expands task environments by composing a finite library of reusable services. Coding agents turn tool specifications into verified services with typed states and shared interfaces, while a world model handles unreliable tools. Verified trajectories support supervised fine-tuning (SFT), while a Completion-Focused Rubric Reward guides reinforcement learning (RL). Using 448 services exposing 10,130 tools, 3K SFT trajectories, and 1K RL tasks to train Qwen3.6-35B-A3B, the method improves performance by 9.17 points on average across eight benchmarks and surpasses frontier models like Claude Opus 4.6 on AutomationBench.
CompoWorld constructs 448 services exposing 10,130 tools and uses 3K SFT trajectories and 1K RL tasks to train Qwen3.6-35B-A3B.
CompoWorld improves on its backbone by 9.17 points on average across eight benchmarks.
On AutomationBench, it surpasses frontier models such as Claude Opus 4.6 and leads all compared agent-specialized 35B-A3B models.