AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs
The study introduces AgentWorld, a benchmark of 100 human-annotated tasks with 100 augmented variants to evaluate long-horizon, multi-agent collaboration spanning over 50 interaction rounds across an MMORPG sandbox. It requires 3-20 agents with asymmetric roles to coordinate via communication, joint planning, and resource sharing in a blackbox setting. To quantify collaboration effectiveness beyond binary success, the authors propose Causal Collaboration Effectiveness (CCE), a graph-based metric tracing causal dependencies between agent actions. Experiments with top models show that even the best model achieves only 52.0% task success, exposing systematic failures in communication and role maintenance.
Even the best model achieves only 52.0% task success on the AgentWorld benchmark.