Skip to content
Launchpad Library logo

All resources / Technology & Ethics

FreeResearch Paper
Technology & Ethics

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

The PDF opens on the site where the paper is published.

Authors
Raphael Shu, Yusen Zhang, Young Min Cho, Jin Mo Yang, Yuan Yuan, Wenliang Zheng, Sharath Chandra Guntuku, Lyle Ungar, Zhou Yu, Rui Zhang
Venue
arXiv (preprint — not peer-reviewed)
Published
September 25, 2026
ID
arXiv:2609.31590

Abstract

Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance, failing to isolate and highlight genuine collaboration capabilities of LLM-based agents. We introduce AgentWorld, a benchmark of 100 human-annotated tasks (with 100 augmented variants) for evaluating long-horizon, multi-agent collaboration. Tasks span 50+ interaction rounds across a rich MMORPG sandbox and require 3-20 agents with asymmetric roles and abilities to coordinate through communication, joint planning, and resource sharing under a blackbox setting where each agent acts independently without access to others' internal states. To quantify collaboration effectiveness in addition to conventional binary task success, we propose Causal Collaboration Effectiveness (CCE), a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome. Experiments with Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, and DeepSeek R1-70B show that even the best model achieves only 52.0% task success, with systematic failure modes including communication breakdowns, role confusion, and inability to maintain shared plans across rounds. AgentWorld is fully open-source.

The authors' own abstract, copied from arxiv.org.

How to cite it

Built from the details listed on this page. Check them against the paper before you submit.

APA
Shu, R., Zhang, Y., Cho, Y. M., Yang, J. M., Yuan, Y., Zheng, W., Guntuku, S. C., Ungar, L., Yu, Z., & Zhang, R. (2026). AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs. arXiv (preprint — not peer-reviewed). https://arxiv.org/abs/2609.31590
MLA
Shu, Raphael, et al. "AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs." arXiv (preprint — not peer-reviewed), 2026, https://arxiv.org/abs/2609.31590.
BibTeX
@article{shu2026agentworld,
  title = {AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs},
  author = {Raphael Shu and Yusen Zhang and Young Min Cho and Jin Mo Yang and Yuan Yuan and Wenliang Zheng and Sharath Chandra Guntuku and Lyle Ungar and Zhou Yu and Rui Zhang},
  journal = {arXiv (preprint — not peer-reviewed)},
  year = {2026},
  url = {https://arxiv.org/abs/2609.31590},
}

What it is

A benchmark of 100 human-annotated tasks (plus 100 augmented variants) for evaluating long-horizon, multi-agent collaboration, designed to isolate genuine collaboration capabilities of LLM-based agents rather than short competitive interactions.

Topics

Added Oct 2, 2026 · 0 opens