CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — evaluation 7 upvotes

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

QUESTION — How can the long-term collaboration capabilities of multiple LLM agents be accurately measured in a complex sandbox environment?

The study introduces AgentWorld, a benchmark of 100 human-annotated tasks with 100 augmented variants to evaluate long-horizon, multi-agent collaboration spanning over 50 interaction rounds across an MMORPG sandbox. It requires 3-20 agents with asymmetric roles to coordinate via communication, joint planning, and resource sharing in a blackbox setting. To quantify collaboration effectiveness beyond binary success, the authors propose Causal Collaboration Effectiveness (CCE), a graph-based metric tracing causal dependencies between agent actions. Experiments with top models show that even the best model achieves only 52.0% task success, exposing systematic failures in communication and role maintenance.

Even the best model achieves only 52.0% task success on the AgentWorld benchmark.

taesiri · 25 Sept 2026 read the original ↗
↑